1396 norm sets
We collected subjective frequency, age-of-acquisition, and imageability norms for 319 acronyms from French adults. Objective printed frequency, bigram frequency, and lengths in letters, phonemes, and syllables, as well as orthographic neighbors, were computed. The time taken to read acronyms aloud was also recorded. Correlational analyses indicated that the relations between the psycholinguistic variables were similar to those usually found for common words (e.g., highly imageable acronyms were more frequent and learned earlier in life than less imageable acronyms), but were generally weaker in the former than in the latter. Linear mixed-model analyses performed on the reading latencies revealed that the main determinants were the voicing feature of initial phonemes, the type of pronunciation of the acronyms (ambiguous vs. unambiguous, typical vs. atypical characteristics), length (number of letters and number of syllables), together with bigram frequency, printed frequency, and imageability. Both objective frequency and imageability interacted reliably with the ambiguous typical and ambiguous atypical properties. Accuracy was predicted by the number of letters and by imageability factors: More errors occurred on longer than on shorter acronyms, and also more errors on less imageable than on more imageable acronyms. The theoretical and methodological implications of the findings for the understanding of acronym reading are discussed. The entire set of norms and the acronym reading times (and accuracy scores), together with the acronym definitions, are provided as supplemental materials.
In order to explore the role of the main psycholinguistic variables on visual word recognition, several mega-studies have been conducted in English in recent years. Nevertheless, because the effects of these variables depend on the regularity of the orthographic system, studies must also be done in other languages with different characteristics. The goal of this work was to conduct a lexical decision study in Spanish, a language with a shallow orthography and a high number of words. The influence of psycholinguistic variables on latencies corresponding to 2,765 words was assessed by means of linear mixed-effects modeling. The results show that some variables, such as frequency or age of acquisition, have significant effects on reaction times regardless of the type of words used. Other variables, such as orthographic neighborhood or imageability, were significant only in specific groups of words. Our results highlight the importance of taking into account the peculiarities of each spelling system in the development of reading models.
We present Chinese translation norms for 1429 English words. Chinese-English bilinguals (N=28) were asked to provide the first Chinese translation that came to mind for 1429 English words. The results revealed that 71{\%} of the English words received more than one correct translation indicating the large amount of translation ambiguity when translating from English to Chinese. The relationship between translation ambiguity and word frequency, concreteness and language proficiency was investigated. Although the significant correlations were not strong, results revealed that English word frequency was positively correlated with the number of alternative translations, whereas English word concreteness was negatively correlated with the number of translations. Importantly, regression analyses showed that the number of Chinese translations was predicted by word frequency and concreteness. Furthermore, an interaction between these predictors revealed that the number of translations was more affected by word frequency for more concrete words than for less concrete words. In addition, mixed-effects modelling showed that word frequency, concreteness and English language proficiency were all significant predictors of whether or not a dominant translation was provided. Finally, correlations between the word frequencies of English words and their Chinese dominant translations were higher for translation-unambiguous pairs than for translation-ambiguous pairs. The translation norms are made available in a database together with lexical information about the words, which will be a useful resource for researchers investigating Chinese-English bilingual language processing.
Timed picture naming was compared in seven languages that vary along dimensions known to affect lexical access. Analyses over items focused on factors that determine cross-language universals and cross-language disparities. With regard to universals, number of alternative names had large effects on reaction time within and across languages after target–name agreement was controlled, suggesting inhibitory effects from lexical competitors. For all the languages, word frequency and goodness of depiction had large effects, but objective picture complexity did not. Effects of word structure variables (length, syllable structure, compounding, and initial frication) varied markedly over languages. Strong cross-language correlations were found in naming latencies, frequency, and length. Other-language frequency effects were observed (e.g., Chinese frequencies predicting Spanish reaction times) even after within-language effects were controlled (e.g., Spanish frequencies predicting Spanish reaction times). These surprising cross-language correlations challenge widely held assumptions about the lexical locus of length and frequency effects, suggesting instead that they may (at least in part) reflect familiarity and accessibility at a conceptual level that is shared over languages.
? 2016 Psychonomic Society, Inc.Semantic feature production norms provide many quantitative measures of different feature and concept variables that are necessary to solve some debates surrounding the nature of the organization, both normal and pathological, of semantic memory. Despite the current existence of norms for different languages, there are still no published norms in Spanish. This article presents a new set of norms collected from 810 participants for 400 living and nonliving concepts among Spanish speakers. These norms consist of empirical collections of features that participants used to describe the concepts. Four files were elaborated: a concept?feature file, a concept?concept matrix, a feature?feature matrix, and a significantly correlated features file. We expect that these norms will be useful for researchers in the fields of experimental psychology, neuropsychology, and psycholinguistics.
The present study provides Italian normative measures for 266 line drawings belonging to the new set of pictures developed by Lotto, Dell'Acqua, and Job (in press). The pictures have been standardized on the following measures: number of letters, number of syllables, name frequency, within-category typicality, familiarity, age of acquisition, name agreement, and naming time. In addition to providing the measures, the present study focuses on indirect and direct comparisons (i.e., correlations) of the present norms with databases provided by comparable studies in Italian (in which normative data were collected with Snodgrass {\&} Vanderwart's set of pictures; Nisi, Longoni, {\&} Snodgrass, 2000), in British English (Barry, Morrison, {\&} Ellis, 1997), in American English (Snodgrass {\&} Vanderwart, 1980; Snodgrass {\&} Yuditsky, 1996), in French (Alario {\&} Ferrand, 1999), and in Spanish (Sanfeliu {\&} Fernandez, 1996)
Most current models of research on emotion recognize valence (how pleasant a stimulus is) and arousal (the level of activation or intensity that a stimulus elicits) as important components in the classification of affective experiences (Barrett, 1998; Kuppens, Tuerlinckx, Russell, {\&} Barrett, 2012). Here we present a set of norms for valence and arousal for a very large set of Spanish words, including items from a variety of frequencies, semantic categories, and parts of speech, including a subset of conjugated verbs. In this regard, we found that there were significant but very small differences between the ratings for conjugations of the same verb, validating the practice of applying the ratings for infinitives to all derived forms of the verb. Our norms show a high degree of reliability and are strongly correlated with those of Redondo, Fraga, Padr{\'{o}}n, and Comesa{\~{n}}a's (2007) Spanish version of the influential Affective Norms for English Words (Bradley {\&} Lang, 1999), as well as those from Warriner, Kuperman, and Brysbaert (2013), the largest available set of emotional norms for English words. Additionally, we included measures of word prevalence-that is, the percentage of participants that knew a particular word-for each variable (Keuleers, Stevens, Mandera, {\&} Brysbaert, 2015). Our large set of norms in Spanish not only will facilitate the creation of stimuli and the analysis of texts in that language, but also will be useful for cross-language comparisons and research on emotional aspects of bilingualism. The norms can be downloaded and available as a supplementary materials to this article.
This study presents a normative database of Spanish restricted length word stems that provides useful information for the selection of stimuli in memory experiments with Word Stem Completion (WSC) tasks. The database includes indices relative to stems (total baseline completion, priming baseline completion, priming, number of completions, ratio between given and deleted letters, and syllabic structure), and indices relative to characteristics of the words used to obtain the stems (frequency, familiarity, number of meanings, length, number of syllables, arousal, and valence). A WSC task was performed by 515 participants to calculate priming and baseline indices. An Exploratory Factor Analysis showed that these indices are grouped in four factors: perceptual, lexical, emotional, and response competition. Stepwise regression analyses performed with these factors showed that the lexical, response competition, and perceptual factors predict priming baseline completion, while only the lexical factor predicts priming. The model that best explains the relationship between priming and priming baseline completion was a cubic model, and the optimum baseline values for achieving priming were between .31 and .36. These norms can be downloaded as Supplemental Materials for this article from https://nuvol.uv.es/owncloud/index.php/s/hpj9by1qbENdjfj .
textcopyright} 2016, The Author(s). In this article, we introduce HelexKids, an online written-word database for Greek-speaking children in primary education (Grades 1 to 6). The database is organized on a grade-by-grade basis, and on a cumulative basis by combining Grade 1 with Grades 2 to 6. It provides values for Zipf, frequency per million, dispersion, estimated word frequency per million, standard word frequency, contextual diversity, orthographic Levenshtein distance, and lemma frequency. These values are derived from 116 textbooks used in primary education in Greece and Cyprus, producing a total of 68,692 different word types. HelexKids was developed to assist researchers in studying language development, educators in selecting age-appropriate items for teaching, as well as writers and authors of educational books for Greek/Cypriot children. The database is open access and can be searched online at www.helexkids.org.
Words are widely used as stimuli in cognitive research. Because of their complexity, using words requires strict control of their objective (lexical and sublexical) and subjective properties. In this work, we present the Minho Word Pool (MWP), a dataset that provides normative values of imageability, concreteness, and subjective frequency for 3,800 (European) Portuguese words-three subjective measures that, in spite of being used extensively in research, have been scarce for Portuguese. Data were collected with 2,357 college students who were native speakers of European Portuguese. The participants rated 100 words drawn randomly from the full set for each of the three subjective indices, using a Web survey procedure (via a URL link). Analyses comparing the MWP ratings with those obtained for the same words from other national and international databases showed that the MWP norms are reliable and valid, thus providing researchers with a useful tool to support research in all neuroscientific areas using verbal stimuli. The MWP norms can be downloaded along with this article or from http://p-pal.di.uminho.pt/about/databases .
Color has the ability to influence a variety of human behaviors, such as object recognition, the identification of facial expressions, and the ability to categorize stimuli as positive or negative. Researchers have started to examine the relationship between emotional words and colors, and the findings have revealed that brightness is often associated with positive emotional words and darkness with negative emotional words (e.g., Meier, Robinson, {\&} Clore, Psychological Science, 15, 82-87, 2004). In addition, words such as anger and failure seem to be inherently associated with the color red (e.g., Kuhbandner {\&} Pekrun). The purpose of the present study was to construct norms for positive and negative emotion and emotion-laden words and their color associations. Participants were asked to provide the first color that came to mind for a set of 160 emotional items. The results revealed that the color RED was most commonly associated with negative emotion and emotion-laden words, whereas YELLOW and WHITE were associated with positive emotion and emotion-laden words, respectively. The present work provides researchers with a large database to aid in stimulus construction and selection.
Textual analysis has been applied to various fields, such as discourse analysis, corpus studies, text leveling, and automated essay evaluation. Several tools have been devel-oped for analyzing texts written in alphabetic languages such as English and Spanish. However, currently there is no tool available for analyzing Chinese-language texts. This article introduces a tool for the automated analysis of simplified and traditional Chinese texts, called the Chinese Readability Index Explorer (CRIE). Composed of four subsystems and incorporating 82 multilevel linguistic features, CRIE is able to conduct the major tasks of segmentation, syntactic parsing, and feature extraction. Furthermore, the integration of linguis-tic features with machine learning models enables CRIE to provide leveling and diagnostic information for texts in lan-guage arts, texts for learning Chinese as a foreign language, and texts with domain knowledge. The usage and validation of the functions provided by CRIE are also introduced.
Written symbols such as letters have been extensively used in cognitive psychology, be it to understand their contribution to written word recognition or to examine processes involved in other mental functions. Sometimes, however, researchers want to manipulate letters while removing their associated characteristics. A powerful solution to do so is to use new characters, devised to be highly similar to letters, but without associated sound or name. Given the growing use of artificial characters in experimental paradigms, the aim of the present study was to make available the Brussels Artificial Character Sets (BACS), two full, strictly controlled, and portable sets of artificial characters for a broad range of experimental situations.
Most experimental research making use of the Japanese language has involved the 1945 officially standardized kanji (Japanese logographic characters) in the Jōyō kanji list (originally announced by the Japanese government in 1981). However, this list was extensively modified in 2010: five kanji were removed and 196 kanji were added; the latest revision of the list now has a total of 2136 kanji. Using an up-to-date corpus consisting of 11 years' worth of articles printed in the Mainichi Newspaper (2000-2010), we have constructed two novel databases that can be used in psychological research using the Japanese language: (1) a database containing a wide variety of properties on the latest 2136 Jōyō kanji, and (2) a novel database containing 27,950 two-kanji compound words (or jukugo). Based on these two databases, we have created an interactive website ( www.kanjidatabase.com ) to retrieve and store linguistic information to be used in psychological and linguistic experiments. The present paper reports the most important characteristics for the new databases, as well as their value for experimental psychological and linguistic research
Emotion expression in human-human interaction takes place via various types of information, including body motion. Research on the perceptual-cognitive mechanisms underlying the processing of natural emotional body language can benefit greatly from datasets of natural emotional body expressions that facilitate stimulus manipulation and analysis. The existing databases have so far focused on few emotion categories which display predominantly prototypical, exaggerated emotion expressions. Moreover, many of these databases consist of video recordings which limit the ability to manipulate and analyse the physical properties of these stimuli. We present a new database consisting of a large set (over 1400) of natural emotional body expressions typical of monologues. To achieve close-to-natural emotional body expressions, amateur actors were narrating coherent stories while their body movements were recorded with motion capture technology. The resulting 3-dimensional motion data recorded at a high frame rate (120 frames per second) provides fine-grained information about body movements and allows the manipulation of movement on a body joint basis. For each expression it gives the positions and orientations in space of 23 body joints for every frame. We report the results of physical motion properties analysis and of an emotion categorisation study. The reactions of observers from the emotion categorisation study are included in the database. Moreover, we recorded the intended emotion expression for each motion sequence from the actor to allow for investigations regarding the link between intended and perceived emotions. The motion sequences along with the accompanying information are made available in a searchable MPI Emotional Body Expression Database. We hope that this database will enable researchers to study expression and perception of naturally occurring emotional body expressions in greater depth.
Words are frequently used as stimuli in cognitive psychology experiments, for example, in recognition memory studies. In these experiments, it is often desirable to control for the words' psycholinguistic properties because differences in such properties across experi-mental conditions might introduce undesirable confounds. In order to avoid confounds, stud-ies typically check to see if various affective and lexico-semantic properties are matched across experimental conditions, and so databases that contain values for these properties are needed. While word ratings for these variables exist in English and other European lan-guages, ratings for Chinese words are not comprehensive. In particular, while ratings for sin-gle characters exist, ratings for two-character words—which often have different meanings than their constituent characters, are scarce. In this study, ratings for 292 two-character Chi-nese nouns were obtained from Cantonese speakers in Hong Kong. Affective variables, including valence and arousal, and lexico-semantic variables, including familiarity, concrete-ness, and imageability, were rated in the study. The words were selected from a film subtitle database containing word frequency information that could be extracted and listed along-side the resulting ratings. Overall, the subjective ratings showed good reliability across all rated dimensions, as well as good reliability within and between the different groups of par-ticipants who each rated a subset of the words. Moreover, several well-established relation-ships between the variables found consistently in other languages were also observed in this study, demonstrating that the ratings are valid. The resulting word database can be used in studies where control for the above psycholinguistic variables is critical to the research design. PLOS ONE | https://doi.org/10.1371/journal.pone.0174569 March 27, 2017 1 / 16 a1111111111 a1111111111 a1111111111 a1111111111 a1111111111 OPEN ACCESS Citation: Yee LTS (2017) Valence, arousal, familiarity, concreteness, and imageability ratings for 292 two-character Chinese nouns in Cantonese speakers in Hong Kong. PLoS ONE 12(3):
In many research domains, researchers have employed gradually morphing pictures to study perception under ambiguity. Despite their inherent utility, only a limited number of stimulus sets are available, and those sets vary substantially in quality and perceptual complexity. Here we present normative data for 40 morphing picture series. In all sets, line drawings of pictures of common objects are morphed over 15 iterations into a completely different object. Objects are either morphed from an animate to an inanimate object (or vice versa) or morphed within the animate and inanimate object categories. These pictures, together with the normative naming data presented here, will be of value for research on a diverse range of questions, from perceptual processing to decision making.
In languages where the position of lexical stress within a word is not predictable from print, readers rely on distributional information extracted from the lexicon in order to assign stress. Lexical databases are thus especially important for researchers willing to address stress assignment in those languages. Here we present Q2Stress, a new database aimed to fill the lack of such a resource for Italian. Q2Stress includes multiple cues readers may use in assigning stress, such as type and token frequency of stress patterns as well as their distribution with respect to number of syllables, grammatical category, word beginnings, word endings, and consonant-vowel structures. Furthermore, for the first time, data for both adults and children are available. Q2Stress may help researchers to answer empirical as well as theoretical questions about stress assignment and stress-related issues, and more in general, to explore the orthography-to-phonology relation in reading. Q2Stress is designed as a user-friendly resource, as it comes with scripts allowing researchers to explore and select their own stimuli according to several criteria as well as summary tables for overall data analysis.
Perceptual information is important for the meaning of nouns. We present modality exclusivity norms for 485 Dutch nouns rated on visual, auditory, haptic, gustatory, and olfactory associations. We found these nouns are highly multimodal. They were rated most dominant in vision, and least in olfaction. A factor analysis identified two main dimensions: one loaded strongly on olfaction and gustation (reflecting joint involvement in flavor), and a second loaded strongly on vision and touch (reflecting joint involvement in manipulable objects). In a second study, we validated the ratings with similarity judgments. As expected, words from the same dominant modality were rated more similar than words from different dominant modalities; but – more importantly – this effect was enhanced when word pairs had high modality strength ratings. We further demonstrated the utility of our ratings by investigating whether perceptual modalities are differentially experienced in space, in a third study. Nouns were categorized into their dominant modality and used in a lexical decision experiment where the spatial position of words was either in proximal or distal space. We found words dominant in olfaction were processed faster in proximal than distal space compared to the other modalities, suggesting olfactory information is mentally simulated as “close” to the body. Finally, we collected ratings of emotion (valence, dominance, and arousal) to assess its role in perceptual space simulation, but the valence did not explain the data. So, words are processed differently depending on their perceptual associations, and strength of association is captured by modality exclusivity ratings.
Faces impart exhaustive information about their bearers, and are widely used as stimuli in psychological research. Yet many extant facial stimulus sets have sub- stantially less detail than faces encountered in real life. In this paper, we describe a new database of facial stimuli, the Multi-Racial Mega-Resolution database (MR2). The MR2 includes 74 extremely high resolution images of European, African, and East Asian faces. This database provides a high-quality, diverse, naturalistic, and well-controlled facial image set for use in research. The MR2 is available under a Creative Commons license, and may be accessed online.
Concreteness ratings are presented for 37,058 English words and 2,896 two-word expressions (such as zebra crossing and zoom in), obtained from over 4,000 participants by means of a norming study using Internet crowdsourcing for data collection. Although the instructions stressed that the assessment of word concreteness would be based on experiences involving all senses and motor responses, a comparison with the existing concreteness norms indicates that participants, as before, largely focused on visual and haptic experiences. The reported data set is a subset of a comprehensive list of English lemmas and contains all lemmas known by at least 85 {\%} of the raters. It can be used in future research as a reference list of generally known English lemmas.
We have developed and tested 144 compound remote associate problems. Across eight experiments, 289 participants were given four time limits (2 sec, 7 sec, 15 sec, or 30 sec) for solving each problem. This paper provides a brief overview of the problems and normative data regarding the percentage of participants solving, and mean time-to-solution for, each problem at each time limit. These normative data can be used in selecting problems on the basis of difficulty or mean time necessary for reaching a solution.
In the present study, we present normative ratings of free association for 139 European Portuguese (EP) words among 7- to 8-, 9- to 10-, and 11- to 12-year-old children attending the 3rd, 5th, and 7th grades of elementary and middle school in Portugal. For each word, five indices are presented: (a) the percentage of associates, (b) the strength of the first associate, (c) the strength of the second associate, (d) the distance between the first and second associates, and (e) the percentage of idiosyncratic responses. Additionally, grade-level frequency values for each word from the ESCOLEX database (Soares et al., in press) are also provided. As expected, the results revealed developmental changes in the knowledge organization of the children, which occurred at the ages of 9–10 (5th grade) and remained stable in the 11- to 12-year-old children (7th grade). Specifically, we observed a decrease in the percentages of associates and idiosyncratic responses, as well as an increase in the strengths of the first and second associates from the 3rd to the 5th grade. Moreover, a comparative analysis with the previous work of Carneiro, Albuquerque, Fernandez, and Esteves (2004) on EP and Macizo, G{\'{o}}mez-Ariza, and Bajo (2000) on Spanish, for the subsets of common words (16 and 58, respectively), showed that the present norms fit well with previous EP data, but differ from the Spanish data.
textcopyright} 2016 Psychonomic Society, Inc.Using a megastudy approach, we developed a database of lexical variables and lexical decision reaction times and accuracy rates for more than 25,000 traditional Chinese two-character compound words. Each word was responded to by about 33 native Cantonese speakers in Hong Kong. This resource provides a valuable adjunct to influential mega-databases, such as the Chinese single-character, English, French, and Dutch Lexicon Projects. Three analyses were conducted to illustrate the potential uses of the database. First, we compared the proportion of variance in lexical decision performance accounted for by six word frequency measures and established that the best predictor was Cai and Brysbaert's (PLoS One, 5, e10729, 2010) contextual diversity subtitle frequency. Second, we ran virtual replications of three previously published lexical decision experiments and found convergence between the original experiments and the present megastudy. Finally, we conducted item-level regression analyses to examine the effects of theoretically important lexical variables in our normative data. This is the first publicly available large-scale repository of behavioral responses pertaining to Chinese two-character compound word processing, which should be of substantial interest to psychologists, linguists, and other researchers.
In this study, we present the normative values of the adaptation of the International Affective Digitized Sounds (IADS-2; Bradley {\&} Lang, 2007a) for European Portuguese (EP). The IADS-2 is a standardized database of 167 naturally occurring sounds that is widely used in the study of emotions. The sounds were rated by 300 college students who were native speakers of EP, in the three affective dimensions of valence, arousal, and dominance, by using the Self-Assessment Manikin (SAM). The aims of this adaptation were threefold: (1) to provide researchers with standardized and normatively rated affective sounds to be used with an EP population; (2) to investigate sex and cultural differences in the ratings of affective dimensions of auditory stimuli between EP and the American (Bradley {\&} Lang, 2007a) and Spanish (Fern{\'{a}}ndez-Abascal et al., Psicothema 20:104-113 2008; Redondo, Fraga, Padr{\'{o}}n, {\&} Pi{\~{n}}eiro, Behavior Research Methods 40:784-790 2008) standardizations; and (3) to promote research on auditory affective processing in Portugal. Our results indicated that the IADS-2 is a valid and useful database of digitized sounds for the study of emotions in a Portuguese context, allowing for comparisons of its results with those of other international studies that have used the same database for stimulus selection. The normative values of the EP adaptation of the IADS-2 database can be downloaded along with the online version of this article.
In the present study, we collected valence, arousal, concreteness, familiarity, imageability, and context availabili-ty ratings for a total of 1,100 Chinese words. The ratings for all variables were collected with 9-point Likert scales. We tested the reliability of the present database by comparing it to the extant Chinese Affective Word System, and performed split-half correlations for all six variables. We then evaluated the relationships between all variables. Regarding the affective variables, we found a typical quadratic relation between va-lence and arousal, in line with previous findings. Likewise, significant correlations were found between the semantic var-iables. Importantly, we explored the relationships between ratings for the affective variables (i.e., valence and arousal) and concreteness ratings, suggesting that valence and arousal ratings can predict concreteness ratings. This database of af-fective norms will be a valuable source of information for emotion research that makes use of Chinese words, and will enable researchers to use highly controlled Chinese verbal stimuli to more reliably investigate the relation between cog-nition and emotion.
The Battig and Montague (1969) category norms have been an invaluable tool for researchers in many fields, with a recent literature search revealing their use in over 1600 projects published in more than 200 different journals. Since 1969, numerous changes have occurred culturally that warrant the collection of new normative data. For instance, in the mid-1960s, the waltz was a popular dance, and undergraduates wore rubbers on their feet. To meet the need for updated norms, we report an expanded version of the Battig and Montague (1969) norms, based on responses from three different sites varying in geographical locations within the United States. The norms were expanded to include new categories (e.g., ad hoc categories) and new measures, most notably latencies for the generated responses. Analyses demonstrated high levels of geographical stability across the new sites, with lower and more variable levels of generational stability between the Battig and Montague norms and the current norms.
This article describes a new software tool called RadicalLocator that can be used to automatically identify (e.g., for visual inspection) individual target radicals (i.e., groups of strokes) in written Chinese characters. We first briefly clarify why this software is useful for research purposes and discuss the factors that make this pattern recognition task so difficult. We then describe how the software can be downloaded and installed, and used to identify the radicals in characters for the purposes of, for example, selecting materials for psycholinguistic experiments. Finally, we discuss several known limitations of the software and heuristics for addressing them.
The Corpus of Contemporary American English ( COCA ), which was released online in early 2008, is the first large and diverse corpus of American English. In this paper, we first discuss the design of the corpus — which contains more than 385 million words from 1990–2008 (20 million words each year), balanced between spoken, fiction, popular magazines, newspapers, and academic journals. We also discuss the unique relational databases architecture, which allows for a wide range of queries that are not available (or are quite difficult) with other architectures and interfaces. To conclude, we consider insights from the corpus on a number of cases of genre-based variation and recent linguistic variation, including an extended analysis of phrasal verbs in contemporary American English.
This paper describes a computerised database of psycholinguistic information. Semantic, syntactic, phonological and orthographic information about some or all of the 98,538 words in the database is accessible, by using a specially-written and very simple programming language. Word-association data are also included in the database. Some examples are given of the use of the database for selection of stimuli to be used in psycholinguistic experimentation or linguistic research.
In the present study, we report naming latencies and norms for 327 photos of objects in Dutch. We provide norms for eight psycholinguistic variables: age of acquisition, familiarity, imageability, image agreement, objective and subjective visual complexity, word frequency, word length in syllables and letters, and name agreement. Furthermore, multiple regression analyses revealed that the significant predictors of photo-naming latencies were name agreement, word frequency, imageability, and image agreement. The naming latencies, norms, and stimuli are provided as supplemental materials.
Druks and Masterson [Druks J, Masterson J. An object and action naming battery with pairwise matching on various psycholinguistic characteristics. (Submitted).] produced a set of 164 object and 102 action pictures which are matched on a range of variables known to affect the availability of picture names. In the present paper we describe the development of the set of pictures and provide the verbal labels for the pictures together with their printed word frequency values and ratings for age-of-acquisition, familiarity and imageability; we also present semantic categories for the verbal labels. Finally, we give visual complexity ratings for the pictures. The materials can be used for a range of psycholinguistic experiments and also for assessment and remediation with clinical populations.
NIM is Web-based software developed to help experimenters with some of the usual tasks carried out in psycholinguistic studies. It allows the user to search for words according to several variables, such as length, matching substrings, lexical frequency, or part of speech, in English, Spanish, and Catalan. NIM also provides the user with the possibilities to obtain different word metrics, such as lexical frequency, length, and part of speech; to find intralanguage and cross-language lexical neighbors; and to get control words for critical stimuli. Regardless of the language used, the program also enables the user to get the orthographic similarity between word pairs and to identify repeated items in lists of experimental stimuli. NIM is free and is publicly available at http://psico.fcep.urv.cat/utilitats/nim/ .
We review recent evidence indicating that researchers in experimental psychology may have used suboptimal estimates of word frequency. Word frequency measures should be based on a corpus of at least 20 million words that contains language participants in psychology experiments are likely to have been exposed to. In addition, the quality of word frequency measures should be ascertained by correlating them with behavioral word processing data. When we apply these criteria to the word frequency measures available for the German language, we find that the commonly used Celex frequencies are the least powerful to predict lexical decision times. Better results are obtained with the Leipzig frequencies, the dlexDB frequencies, and the Google Books 2000–2009 frequencies. However, as in other languages the best performance is observed with subtitle-based word frequencies. The SUBTLEX-DE word frequencies collected for the present ms are made available in easy-to-use files and are free for educational purposes.
We present word frequencies based on subtitles of British television programmes. We show that the SUBTLEX-UK word frequencies explain more of the variance in the lexical decision times of the British Lexicon Project than the word frequencies based on the British National Corpus and the SUBTLEX-US frequencies. In addition to the word form frequencies, we also present measures of contextual diversity part-of-speech specific word frequencies, word frequencies in children programmes, and word bigram frequencies, giving researchers of British English access to the full range of norms recently made available for other languages. Finally, we introduce a new measure of word frequency, the Zipf scale, which we hope will stop the current misunderstandings of the word frequency effect.
Measures of icon designs rely heavily on surveys of the perceptions of population samples. Thus, measuring the extent to which changes in the structure of an icon will alter its perceived complexity can be costly and slow. An automated system capable of producing reliable estimates of perceived complexity could reduce development costs and time. Measures of icon complexity developed by Garcia, Badre, and Stasko (1994) and McDougall, Curry, and de Bruijn (1999) were correlated with six icon properties measured using Matlab (MathWorks, 2001) software, which uses image-processing techniques to measure icon properties. The six icon properties measured were icon foreground, the number of objects in an icon, the number of holes in those objects, and two calculations of icon edges and homogeneity in icon structure. The strongest correlates with human judgments of perceived icon complexity (McDougall et al., 1999) were structural variability (r(s) = .65) and edge information (r(s) = .64).
In this article, we present the first open-access lexical database that provides phonological representations for 120,000 Italian word forms. Each of these also includes syllable boundaries and stress markings and a comprehensive range of lexical statistics. Using data derived from this lexicon, we have also generated a set of derived databases and provided estimates of positional frequency use for Italian phonemes, syllables, syllable onsets and codas, and character and phoneme bigrams. These databases are freely available from phonitalia.org. This article describes the methods, content, and summarizing statistics for these databases. In a first application of this database, we also demonstrate how the distribution of phonological substitution errors made by Italian aphasic patients is related to phoneme frequency.
We developed affective norms for 1,121 Italian words in order to provide researchers with a highly controlled tool for the study of verbal processing. This database was developed from translations of the 1,034 English words present in the Affective Norms for English Words (ANEW; Bradley {\&} Lang, 1999) and from words taken from Italian semantic norms (Montefinese, Ambrosini, Fairfield, {\&} Mammarella, Behavior Research Methods, 45, 440–461, 2013). Participants evaluated valence, arousal, and dominance using the Self-Assessment Manikin (SAM) in a Web survey procedure. Participants also provided evaluations of three subjective psycholinguistic indexes (familiarity, imageability, and concreteness), and five objective psycholinguistic indexes (e.g., word frequency) were also included in the resulting database in order to further characterize the Italian words. We obtained a typical quadratic relation between valence and arousal, in line with previous findings. We also tested the reliability of the present ANEW adaptation for Italian by comparing it to previous affective databases and performing split-half correlations for each variable. We found high split-half correlations within our sample and high correlations between our ratings and those of previous studies, confirming the validity of the adaptation of ANEW for Italian. This database of affective norms provides a tool for future research about the effects of emotion on human cognition.
Formal and semantic overlap across languages plays an important role in bilingual language processing systems. In the present study, Japanese (first language; L1)–English (second language; L2) bilinguals rated 193 Japanese–English word pairs, including cognates and noncognates, in terms of phonological and semantic similarity. We show that the degree of cross-linguistic overlap varies, such that words can be more or less “cognate,” in terms of their phonological and semantic overlap. Bilinguals also translated these words in both directions (L1–L2 and L2–L1), providing a measure of translation equivalency. Notably, we reveal for the first time that Japanese–English cognates are “special,” in the sense that they are usually translated using one English term (e.g., コール /kooru/ is always translated as “call”), but the English word is translated into a greater variety of Japanese words. This difference in translation equivalency likely extends to other nonetymologically related, different-script languages in which cognates are all loanwords (e.g., Korean–English). Norming data were also collected for L1 age of acquisition, L1 concreteness, and L2 familiarity, because such information had been unavailable for the item set. Additional information on L1/L2 word frequency, L1/L2 number of senses, and L1/L2 word length and number of syllables is also provided. Finally, correlations and characteristics of the cognate and noncognate items are detailed, so as to provide a complete overview of the lexical and semantic characteristics of the stimuli. This creates a comprehensive bilingual data set for these different-script languages and should be of use in bilingual word recognition and spoken language research.
Subjective estimations of age of acquisition (AoA) for a large pool of Spanish words were collected from college students in Spain. The average score for each word (based on 50 individual responses, on a scale from 1 to 11) was taken as an AoA indicator, and normative values for a total of 7,039 single words are provided as supplemental materials. Beyond its intrinsic value as a standalone corpus, the largest of its kind for Spanish, the value of the database is enhanced by the fact that it contains most of the words that are currently included in other normative studies, allowing for a more complete characterization of the lexical stimuli that are usually employed in studies with Spanish-speaking participants. The norms are available for downloading as supplemental materials with this article.
Through this study, we aimed to validate a new tool for inducing moods in experimental contexts. Five audio stories with sad, joyful, frightening, erotic, or neutral content were presented to 60 participants (33 women, 27 men) in a within-subjects design, each for about 10 min. Participants were asked (1) to report their moods before and after listening to each story, (2) to assess the emotional content of the excerpts on various emotional scales, and (3) to rate their level of projection into the stories. The results confirmed our a priori emotional classification. The emotional stories were effective in inducing the desired mood, with no difference found between male and female participants. These stories therefore constitute a valuable corpus for inducing moods in French-speaking participants, and they are made freely available for use in scientific research.
Theories of the representation and processing of concepts have been greatly enhanced by models based on information available in semantic property norms. This information relates both to the identity of the features produced in the norms and to their statistical properties. In this article, we introduce a new and large set of property norms that are designed to be a more flexible tool to meet the demands of many different disciplines interested in conceptual knowledge representation, from cognitive psychology to computational linguistics. As well as providing all features listed by 2 or more participants, we also show the considerable linguistic variation that underlies each normalized feature label and the number of participants who generated each variant. Our norms are highly comparable with the largest extant set (McRae, Cree, Seidenberg, {\&} McNorgan, 2005) in terms of the number and distribution of features. In addition, we show how the norms give rise to a coherent category structure. We provide these norms in the hope that the greater detail available in the Centre for Speech, Language and the Brain norms should further promote the development of models of conceptual knowledge. The norms can be downloaded at www.csl.psychol.cam.ac.uk/propertynorms.
Previous evidence has shown that word frequencies calculated from corpora based on film and television subtitles can readily account for reading performance, since the language used in subtitles greatly approximates everyday language. The present study examines this issue in a society with increased exposure to subtitle reading. We compiled SUBTLEX-GR, a subtitled-based corpus consisting of more than 27 million Modern Greek words, and tested to what extent subtitle-based frequency estimates and those taken from a written corpus of Modern Greek account for the lexical decision performance of young Greek adults who are exposed to subtitle reading on a daily basis. Results showed that SUBTLEX-GR frequency estimates effectively accounted for participants' reading performance in two different visual word recognition experiments. More importantly, different analyses showed that frequencies estimated from a subtitle corpus explained the obtained results significantly better than traditional frequencies derived from written corpora.
Virtually no valid materials are available to evaluate confrontation naming in Spanish-English bilingual adults in the U.S. In a recent study, a large group of young Spanish-English bilingual adults were evaluated on An Object and Action Naming Battery (Edmonds {\&} Donovan in Journal of Speech, Language, and Hearing Research 55:359-381, 2012). Rasch analyses of the responses resulted in evidence for the content and construct validity of the retained items. However, the scope of that study did not allow for extensive examination of individual item characteristics, group analyses of participants, or the provision of testing and scoring materials or raw data, thereby limiting the ability of researchers to administer the test to Spanish-English bilinguals and to score the items with confidence. In this study, we present the in-depth information described above on the basis of further analyses, including (1) online searchable spreadsheets with extensive empirical (e.g., accuracy and name agreeability) and psycholinguistic item statistics; (2) answer sheets and instructions for scoring and interpreting the responses to the Rasch items; (3) tables of alternative correct responses for English and Spanish; (4) ability strata determined for all naming conditions (English and Spanish nouns and verbs); and (5) comparisons of accuracy across proficiency groups (i.e., Spanish dominant, English dominant, and balanced). These data indicate that the Rasch items from An Object and Action Naming Battery are valid and sensitive for the evaluation of naming in young Spanish-English bilingual adults. Additional information based on participant responses for all of the items on the battery can provide researchers with valuable information to aid in stimulus development and response interpretation for experimental studies in this population.
The processing of human and nonhuman concepts (e.g., agreeable vs. edible) during basic comprehension and reasoning tasks has become a major topic of scientific inquiry. To ensure that the experimental effects obtained from such studies reflect the hypothesised semantic distinction, potential confounds such as psycholinguistic and/or lexical properties of the exact stimuli chosen need to be addressed. In the current study, normative data of such properties were obtained for a series of 875 French adjectives by asking 8 groups of 20 participants to each rate all words on one dimension of theoretical interest. The collected ratings indicate the extent to which each adjective evokes a sensory experience (concreteness), captures an enduring attribute (temporal stability), refers to a visible characteristic (visibility), denotes a neutral or an affectively laden concept (valence), signifies an attribute of low or high intensity, is familiar to the reader and can be used to describe people and/or inanimate entities such as objects. In addition, for each item its exact grammatical class (adjective vs. past participle adjective), length (i.e., number of letters, number of syllables), and word frequency was retrieved from the lexique3 corpus. The resulting database enables researchers to consider pivotal psycholinguistic and lexical properties when selecting human and nonhuman stimuli for future research.
The Nencki Affective Picture System (NAPS; Marchewka, {\.{Z}}urawski, Jednor{\'{o}}g, {\&} Grabowska, Behavior Research Methods, 2014) is a standardized set of 1,356 realistic, high-quality photographs divided into five categories (people, faces, animals, objects, and landscapes). NAPS has been primarily standardized along the affective dimensions of valence, arousal, and approach-avoidance, yet the characteristics of discrete emotions expressed by the images have not been investigated thus far. The aim of the present study was to collect normative ratings according to categorical models of emotions. A subset of 510 images from the original NAPS set was selected in order to proportionally cover the whole dimensional affective space. Among these, using three available classification methods, we identified images eliciting distinguishable discrete emotions. We introduce the basic-emotion normative ratings for the Nencki Affective Picture System (NAPS BE), which will allow researchers to control and manipulate stimulus properties specifically for their experimental questions of interest. The NAPS BE system is freely accessible to the scientific community for noncommercial use as supplementary materials to this article.
In the present article, we introduce the Nencki Affective Word List (NAWL), created in order to provide researchers with a database of 2,902 Polish words, including nouns, verbs, and adjectives, with ratings of emotional valence, arousal, and imageability. Measures of several objective psycholinguistic features of the words (frequency, grammatical class, and number of letters) are also controlled. The database is a Polish adaptation of the Berlin Affective Word List-Reloaded (BAWL-R; V{\~{o}} et al., Behavior Research Methods 41:534-538, 2009), commonly used to investigate the affective properties of German words. Affective normative ratings were collected from 266 Polish participants (136 women and 130 men). The emotional ratings and psycholinguistic indexes provided by NAWL can be used by researchers to better control the verbal materials they apply and to adjust them to specific experimental questions or issues of interest. The NAWL is freely accessible to the scientific community for noncommercial use as supplementary material to this article.
Research on signed languages offers the opportunity to address many important questions about language that it may not be possible to address via studies of spoken languages alone. Many such studies, however, are inherently limited, because there exist hardly any norms for lexical variables that have appeared to play important roles in spoken language processing. Here, we present a set of norms for age of acquisition, familiarity, and iconicity for 300 British Sign Language (BSL) signs, as rated by deaf signers, in the hope that they may prove useful to other researchers studying BSL and other signed languages. These norms may be downloaded from www.psychonomic.org/archive.
In this article, we present normative data for 2,423 Chinese single-character words. For each word, we report values for the following 15 variables: word frequency, cumulative frequency, homophone density, phonological frequency, age of learning, age of acquisition, number of word formations, number of meanings, number of components, number of strokes, familiarity, concreteness, imageability, regularity, and initial phoneme. To validate the norms, we collected word-naming latencies. Factor analysis and multiple regression analysis show that naming latencies of Chinese single-character words are predicted by frequency, semantics, visual features, and consistency, but not by phonology. These analyses show distinct patterns in word naming between Chinese and alphabetic languages and demonstrate the utility of normative data in the study of nonalphabetic orthographic processing.
Word difficulty varies from language to language; therefore, normative data of verbal stimuli cannot be imported directly from another language. We present mean identification thresholds for the 260 screen-fragmented words corresponding to the total set of Snodgrass and Vanderwart (1980) pictures. Individual words were fragmented in eight levels using Turbo Pascal, and the resulting program was implemented on a PC microcomputer. The words were presented individually to a group of 40 Spanish observers, using a controlled time procedure. An unspecific learning effect was found showing that performance improved due to practice with the task. Finally, of the 11 psycholinguistic variables that previous researchers have shown to affect word identification, only imagery accounted for a significant amount of variance in the threshold values.