1396 norm sets
Word associations to each of the 26 letters of the alphabet were obtained under procedures of single association or continued associations for both upper- and lower-case letters. The results showed a significant relationship between m values and measures of frequency of letters, preferences for letters, and vocal reaction time to letters. The data also showed that m values for each letter were stable within a session. Analyses of the most frequent associations showed a high degree of consistency among the common associations for the single and continued instructional procedures and upper- and lower-case stimulus presentations of the letters. {\textcopyright} 1965 Academic Press Inc. All rights reserved.
The Affective Norms for English Words (ANEW) are a commonly used set of 1,034 words characterized on the affective dimensions of valence, arousal, and dominance. Traditionally, studies of affect have used stimuli characterized along either affective dimensions or discrete emotional categories, but much current research draws on both of these perspectives. As such, stimuli that have been thoroughly characterized according to both of these approaches are exceptionally useful. In an effort to provide researchers with such a characterization of stimuli, we have collected descriptive data on the ANEW to identify which discrete emotions are elicited by each word in the set. Our data, coupled with previous characterizations of the dimensional aspects of these words, will allow researchers to control for or manipulate stimulus properties in accordance with both dimensional and discrete emotional views, and provide an avenue for further integration of these two perspectives. Our data have been archived at www.psychonomic.org/archive/.
- SP{\'{I}}{\v{S}} FONOLOGIE, MOORY ATD$\backslash$r$\backslash$nOn the basis of the lexical corpus created by Amano and Kondo (2000), using the Asahi newspaper, the present study provides frequencies of occurrence for units of Japanese phonemes, morae, and syllables. Among the five vowels, /a/ (23.42{\%}), /i/ (21.54{\%}), /u/ (23.47{\%}), and /o/ (20.63{\%}) showed similar frequency rates, whereas /e/ (10.94{\%}) was less frequent. Among the 12 consonants, /k/ (17.24{\%}), /t/ (15.53{\%}), and /r/ (13.11{\%}) were used often, whereas /p/ (0.60{\%}) and /b/ (2.43{\%}) appeared far less frequently. Among the contracted sounds, /sj/ (36.44{\%}) showed the highest frequency, whereas /mj/ (0.27{\%}) rarely appeared. Among the five long vowels, /aR/ (34.4{\%}) was used most frequently, whereas /uR/ (12.11{\%}) was not used so often. The special sound /N/ appeared very frequently in Japanese. The syllable combination /k/+V+/N/ (19.91{\%}) appeared most frequently among syllabic combinations with the nasal /N/. The geminate (or voiceless obstruent) /Q/, when placed before the four consonants /p/, /t/, /k/, and /s/, appeared 98.87{\%} of the time, but the remaining 1.13{\%} did not follow the definition. The special sounds /R/, /N/, and /Q/ seem to appear very frequently in Japanese, suggesting that they are not special in terms of frequency counts. The present study further calculated frequencies for the 33 newly and officially listed morae/syllables, which are used particularly for describing alphabetic loanwords. In addition, the top 20 bi-mora frequency combinations are reported. Files of frequency indexes may be downloaded from the Psychonomic Society Web archive at http://www.psychonomic.org/archive/.
The present study reports descriptive normative measures for 245 Italian verbal idiomatic expressions. For each of the idiomatic expressions the following variables are reported: Length, Knowledge, Familiarity, Age of Acquisition, Predictability, Syntactic flexibility, Literality and Compositionality. Syntactic flexibility was assessed using five syntactic operations: adverb insertion, adjective insertion, left dislocation, passive and movement. The psycholinguistic relevance of each dimension, their measures and the correlations among them are provided and discussed. The databases are freely available for down-loading from the Psychonomic Society Web archive at www.psychonomic.org/archive/.
In 1981, the Japanese government published a list of the 1,945 basic Japanese kanji (Jooyoo Kanji-hyo), including specifications of pronunciation. This list was established as the standard for kanji usage in print. The database for 1,945 basic Japanese kanji provides 30 cells that explain in detail the various characteristics of kanji. Means, standard deviations, distributions, and information related to previous research concerning these kanji are provided in this paper. The database is saved as a Microsoft Excel 2000 file for Windows. This kanji database is accessible on the Web site of the Oxford Text Archive, Oxford University (http://ota.ahds.ac.uk). Using this database, researchers and educators will be able to conduct planned experiments and organize classroom instruction on the basis of the known characteristics of selected kanji.
The lexical database dlexDB supplies in form of an online database frequency-based norms of numerous process-related word properties for psychological and linguistic research. These values include well known variables such as printed frequency of word form and lemma as documented also in CELEX (Baayen, Piepenbrock und Gulikers, 1995). In addition, we compute new values like frequencies based on syllables, and morphemes as well as frequencies of character chains, and multiple word combinations. The statistics are based on the Kernkorpus des Digitalen Wrterbuchs der deutschen Sprache (DWDS) with over 100 million running words. We illustrate the validity of these norms with new results about fixation durations in sentence reading.
Many cognitive psychological, computational, and neuropsychological approaches to the organisation of semantic memory have incorporated the idea that concepts are, at least partly, represented in terms of their fine-grained features. We asked 20 normal volunteers to provide properties of 64 concrete items, drawn from living and nonliving categories, by completing simple sentence stems (e.g., an owl is {\_}{\_}, has {\_}{\_}, can{\_}{\_}). At a later date, the same participants rated the same concepts for prototypicality and familiarity. The features generated were classified as to type of knowledge (sensory, functional, or encyclopaedic), and also quantified with regard to both dominance (the number of participants specifying that property for that concept) and distinctiveness (the proportion of exemplars within a conceptual category of which that feature was considered characteristic). The results demonstrate that rated prototypicality is related to both the familiarity of the concept and its distance from the average of the exemplars within the same category (the category centroid). The feature database was also used to replicate, resolve, and extend a variety of previous observations on the structure of semantic representations. Specifically, the results of our analyses (1) resolve two conflicting claims regarding the relative ratio of sensory to other kinds of attributes in living vs. nonliving concepts; (2) offer new information regarding the types of features-across different domains-that distinguish concepts from their category coordinates; and (3) corroborate some previous claims of higher intercorrelations between features of living things than those of artefacts.
Semantic differential (SD) factor scores on the Evaluation, Activity, and Potency dimensions are presented for 1,000 most frequently used English words. Also given are the standard errors of the factor scores, the results of several reliability studies, and a listing (for all words) of 3 types of derived scores: polarizations, n Affiliation contents, n Achievement contents. Test-ing procedures and statistics on the sample of raters are detailed. Some uses of the dictionary are suggested, and an example of its use in a study of motivation is presented including empirical results. Conditions favoring further cumulation of SD data are discussed. THE semantic differential (SD) has proven to be an accurate instrument for recording affective associations of stim-uli, particularly to the extent that such as-sociations are culturally or subculturally denned so that measurements may be aver-aged over groups of individuals (Norman, 1959). In a wide variety of studies, includ-ing many involving cross-cultural samples of raters, it has been demonstrated that affective judgments on bipolar adjective scales reliably resolve into three major dimensions or factors which Osgood has named Evaluation, Activity, and Potency 'This paper is part of a doctoral dissertation submitted to the
As researchers explore the complexity of memory and language hierarchies, the need to expand normed stimulus databases is growing. Therefore, we present 1,808 words, paired with their features and concept-concept information, that were collected using previously established norming methods (McRae, Cree, Seidenberg, {\&} McNorgan Behavior Research Methods 37:547-559, 2005). This database supplements existing stimuli and complements the Semantic Priming Project (Hutchison, Balota, Cortese, Neely, Niemeyer, Bengson, {\&} Cohen-Shikora 2010). The data set includes many types of words (including nouns, verbs, adjectives, etc.), expanding on previous collections of nouns and verbs (Vinson {\&} Vigliocco Journal of Neurolinguistics 15:317-351, 2008). We describe the relation between our and other semantic norms, as well as giving a short review of word-pair norms. The stimuli are provided in conjunction with a searchable Web portal that allows researchers to create a set of experimental stimuli without prior programming knowledge. When researchers use this new database in tandem with previous norming efforts, precise stimuli sets can be created for future research endeavors.
We have developed a set of naming and recognition tests for evaluating the retrieval of lexical and conceptual knowledge for actions. As a first step, normative information about 280 items was collected for the following variables: (1) the naming responses elicited by each item, (2) the degree to which the image of each item agreed with a target name, (3) the familiarity to each depicted action, and (4) the visual complexity of each item. This information was used to develop administration and scoring procedures for a standardized test of action naming. The effectiveness and reliability of these procedures were evaluated in a second experiment. In a third experiment, five tests were developed to probe the retrieval of conceptual knowledge: (1) independently of the production of a naming response, (2) in response to pictorial and nonpictorial stimuli, (3) in terms of the attributes associated with specific actions, and (4) in terms of similarities and differences between various actions.
Speeded naming and lexical decision data for 1,661 target words following related and unrelated primes were collected from 768 subjects across four different universities. These behavioral measures have been integrated with demographic information for each subject and descriptive characteristics for every item. Subjects also completed portions of the Woodcock-Johnson reading battery, three attentional control tasks, and a circadian rhythm measure. These data are available at a user-friendly Internet-based repository ( http://spp.montana.edu ). This Web site includes a search engine designed to generate lists of prime-target pairs with specific characteristics (e.g., length, frequency, associative strength, latent semantic similarity, priming effect in standardized and raw reaction times). We illustrate the types of questions that can be addressed via the Semantic Priming Project. These data represent the largest behavioral database on semantic priming and are available to researchers to aid in selecting stimuli, testing theories, and reducing potential confounds in their studies.
Semantic ambiguity is typically measured by sum-ming the number of senses or dictionary definitions that a word has. Such measures are somewhat subjective and may not adequately capture the full extent of variation in word meaning, particularly for polysemous words that can be used in many different ways, with subtle shifts in meaning. Here, we describe an alternative, computationally derived measure of ambiguity based on the proposal that the meanings of words vary continuously as a function of their contexts. On this view, words that appear in a wide range of contexts on diverse topics are more variable in meaning than those that appear in a restricted set of similar contexts. To quantify this variation, we performed latent semantic analysis on a large text corpus to estimate the semantic similarities of different linguistic contexts. From these estimates, we calculated the degree to which the different contexts associated with a given word vary in their meanings. We term this quantity a word's semantic diversity (SemD). We suggest that this approach provides an objective way of quantifying the subtle, context-dependent variations in word meaning that are often present in language. We demonstrate that SemD is correlated with other measures of ambiguity and contextual variability, as well as with frequency and imageability. We also show that SemD is a strong predictor of performance in semantic judgments in healthy individuals and in patients with semantic deficits, accounting for unique variance beyond that of other predictors. SemD values for over 30,000 English words are provided as supplementary materials.
Factors affecting word retrieval were compared in a timed picture-naming paradigm for 520 drawings of objects. In prior timed and untimed studies by Snodgrass
The combining of individual concepts to form an emergent concept is a fundamental aspect of language, yet much less is known about it than about processing isolated words or sentences. To facilitate research on conceptual combination, we provide meaningfulness ratings for a large set of (2,160) noun-noun pairs. Half of these pairs (1,080) are reversed versions of the other half (e.g., SKI JACKET and JACKET SKI), to facilitate the comparison of successful and unsuccessful conceptual combination independently of constituent lexical items. The computer code used for obtaining these ratings through a Web interface is provided. To further enhance the usefulness of this resource, ancillary measures obtained from other sources are also provided for each pair. These measures include associate production norms, contextual relatedness in terms of latent semantic analysis distance, total number of letters, phrase-level usage frequency, and word-level usage frequency summed across the words in each pair. Results of correlation and regression analyses are also provided for a quantitative description of the stimulus set. A subset of these stimuli was used to identify neural correlates of successful conceptual combination Graves, Binder, Desai, Conant, {\&} Seidenberg, (NeuroImage 53:638-646, 2010). The stimuli can be used in other research and also provide benchmark data for evaluating the effectiveness of computational algorithms for predicting meaningfulness of noun-noun pairs.
An operational definition of abstractness in nouns was constructed by using the human discriminative response to identify two points on a scale of abstractness. This scale, consisting of 490 'abstract' and 571 'concrete' nouns, was found to have adequate reliability. When the scale was manipulated as an independent variable, the effect of abstractness on short-term recognition memory was highly significant, 'abstract' nouns being less well remembered than 'concrete' nouns. Frequency was found to be pertinent variable, independent of abstractness, very frequent nouns being less well remembered than some-what rarer nouns.
We make available word-by-word self-paced reading times and eye-tracking data over a sample of English sentences from narrative sources. These data are intended to form a gold standard for the evaluation of computational psycholinguistic models of sentence comprehension in English. We describe stimuli selection and data collection and present descriptive statistics, as well as comparisons between the two sets of reading times.
In order to provide a reliable measure of the similarity of uppercase English letters, a confusion matrix based on 1,200 presentations of each letter was established. To facilitate an analysis of the perceived structural characteristics, the confusion matrix was decomposed according to Luce's choice model into a symmetrical similarity matrix and a response bias vector. The underlying structure of the similarity matrix was assessed with both a hierarchical clustering and a multidimensional scaling procedure. This data is offered to investigators of visual information processing as a valuable tool for controlling not only the overall similarity of the letters in a study, but also their similarity on individual feature dimensions.
A sample of 100 college students ranked the alphabet according to their preference for the appearance of the capital letter. Rankings are presented for the total sample, and for subgroups based on age and sex. Coefficients of concordance among judges are low, but the rankings for the total sample and the age and sex subsamples appear to be quite reliable.
Researchers concerned with the development of cognitive functions are in need of standardized material that can be used with both adults and children. The present article provides normative measures for 400 line drawings viewed by 5- and 6-year-old children. The three variables obtained - name agreement, familiarity, and visual complexity - are important because of their potential effect on memory and other cognitive processes. The normative data collected in the present study indicate that young children are different from adults in both the name most frequently assigned and the number of alternative names provided. The alternative names given by the children are either coordinate names or names of objects that are visually similar to the pictured object. In addition, the failure (to name) rate is higher among young children compared to adults. Thus, we conclude that unequivocal interpretation of age-related differences in cognitive functions can be made only when age-appropriate pictorial stimuli are chosen. {\textcopyright} 1997 Academic Press.
Data from parent reports on 1,803 children--derived from a normative study of the MacArthur Communicative Development Inventories (CDIs)--are used to describe the typical course and the extent of variability in major features of communicative development between 8 and 30 months of age. The two instruments, one designed for 8-16-month-old infants, the other for 16-30-month-old toddlers, are both reliable and valid, confirming the value of parent reports that are based on contemporary behavior and a recognition format. Growth trends are described for children scoring at the 10th-, 25th-, 50th-, 75th-, and 90th-percentile levels on receptive and expressive vocabulary, actions and gestures, and a number of aspects of morphology and syntax. Extensive variability exists in the rate of lexical, gestural, and grammatical development. The wide variability across children in the time of onset and course of acquisition of these skills challenges the meaningfulness of the concept of the modal child. At the same time, moderate to high intercorrelations are found among the different skills both concurrently and predictively (across a 6-month period). Sex differences consistently favor females; however, these are very small, typically accounting for 1{\%}-2{\%} of the variance. The effects of SES and birth order are even smaller within this age range. The inventories offer objective criteria for defining typicality and exceptionality, and their cost effectiveness facilitates the aggregation of large data sets needed to address many issues of contemporary theoretical interest. The present data also offer unusually detailed information on the course of development of individual lexical, gestural, and grammatical items and features. Adaptations of the CDIs to other languages have opened new possibilities for cross-linguistic explorations of sequence, rate, and variability of communicative development.
Normative data on the objective age of acquisition (AoA) for 286 Russian words are presented in this article. In addition, correlations between the objective AoA and subjective ratings, name agreement, picture name agreement, imageability, familiarity, word frequency, and word length are provided, as are correlations between the objective AoA and two measures of exemplar dominance (exemplar generation frequency and the number of times an exemplar was named first). The correlations between the aforementioned variables are generally consistent with the correlations reported in other normative studies. The objective AoA data are highly correlated with the subjective AoA ratings, whereas the correlations between the objective AoA and other psycholinguistic variables are moderate. The correlations between the objective AoA of Russian words and similar data for other languages are moderately high. The complete word norms may be downloaded from supplementary material.
Malay, a language spoken by 250 million people, has a shallow alphabetic orthography, simple syllable structures, and transparent affixation--characteristics that contrast sharply with those of English. In the present article, we first compare the letter-phoneme and letter-syllable ratios for a sample of alphabetic orthographies to highlight the importance of separating language-specific from language-universal reading processes. Then, in order to develop a better understanding of word recognition in orthographies with more consistent mappings to phonology than English, we compiled a database of lexical variables (letter length, syllable length, phoneme length, morpheme length, word frequency, orthographic and phonological neighborhood sizes, and orthographic and phonological Levenshtein distances) for 9,592 Malay words. Separate hierarchical regression analyses for Malay and English revealed how the consistency of orthography-phonology mappings selectively modulates the effects of different lexical variables on lexical decision and speeded pronunciation performance. The database of lexical and behavioral measures for Malay is available at http://brm.psychonomic-journals.org/content/supplemental.
The present study provides Canadian French normative data for 388 line drawings from the European Picture Pool for Oral Naming (Protocole europ{\'{e}}en de d{\'{e}}nomination orale d'images; PEDOI; Kremin et al., 2003). One hundred eighty subjects were equally distributed for age group (18-39,40-59, 60-85), educational level (low, high), and sex. They rated pictures of objects on age of acquisition, name agreement, familiarity, and visual complexity. Syllable length and word frequency were also taken into account. The present study suggests that age of acquisition and name agreement show significant age-related differences. These results show that unequivocal interpretation of age-related differences can be made when age-appropriate norms are used.
Stimulus material for studying object-directed actions is needed in different research contexts, such as action observation, action memory, and imitation. Action items have been generated many times in individual laboratories across the world, but they are used in very few experiments. For future studies in the field, it would be worthwhile to have a larger set of action stimulus material available to a broader research community. Some smaller action databases have already been published, but those often focus on psycholinguistic parameters and static action stimuli. With this article, we introduce an action database with dynamic action stimuli. The database contains action descriptions of 1,754 object-directed actions that have been rated for familiarity in Germany and in China. For 784 of these actions, action video clips are available. With the use of our database, it is possible to identify actions that differ in familiarity between Western and Eastern cultures. This variable may be of interest to some researchers in the field, since it has been shown that familiarity influences action information processing. Action descriptions are listed and categorized in tables that can be downloaded, along with the corresponding video clips, as supplemental material.
We collected number-of-translation norms on 562 Dutch-English translation pairs from several previous studies of cross-language processing. Participants were highly proficient Dutch-English bilinguals. Form and semantic similarity ratings were collected on the 1,003 possible translation pairs. Approximately 40{\%} of the translations were rated as being similar across languages with respect to spelling/sound (i.e., they were cognates). Approximately 45{\%} of the translations were rated as being highly semantically similar across languages. At least 25{\%} of the words in each direction of translation had more than one translation. The form similarity ratings were found to be highly reliable even when obtained with different bilinguals and modified rating procedures. Number of translations and meaning factors significantly predicted the semantic similarity of translation pairs. In future research, these norms may be used to determine the number of translations of words to control for or study this factor. These norms are available at http://www.talkbank.org/norms/tokowicz/.
The aim of the present study was to provide Russian normative data for the Snodgrass and Vanderwart (Behavior Research Methods, Instruments, {\&} Computers, 28, 516-536, 1980) colorized pictures (Rossion {\&} Pourtois, Perception, 33, 217-236, 2004). The pictures were standardized on name agreement, image agreement, conceptual familiarity, imageability, and age of acquisition. Objective word frequency and objective visual complexity measures are also provided for the most common names associated with the pictures. Comparative analyses between our results and the norms obtained in other, similar studies are reported. The Russian norms may be downloaded from the Psychonomic Society supplemental archive.
Ratings of age of acquisition (AoA), imageability, and familiarity were collected for 1,526 words. The methodology made use of a modular approach, in which the full sample of words was divided into five separate blocks. Within each block, each word was rated on each of the three variables by 20 partici- pants (undergraduate students from the University of Bristol). Analyses comparing these ratings to existing norm databases demonstrated that this methodology resulted in high reliability (assessed by Cronbach's ) and validity. The ratings were also transformed to be compatible with the Gilhooly and Logie (1980) norms. This transformation resulted in a set of norms for 3,394 words, which is by far the largest database of ratings for AoA, imageability, and familiarity to date. The resulting database should be useful for researchers interested in manipulating or controlling these factors in word recognition, neuropsychological, or memory studies. These norms can be downloaded from language.psy.bris .ac.uk/bristol{\_}norms.html.
Orthographic transparency metrics for opaque or deep languages, such as French and English, have tended to focus on feedforward and/or feedback directions, with claims made for the influence of both on reading. In the present study, data for five transparency metrics for southern British English, three of which are neither feedforward nor feedback, are presented, demonstrating the complex relationships between the metrics and offering an explanation for feedback effects in children's reading accuracy. The structure of such metrics from a variety of corpus sizes and origins is investigated, and it is concluded that large corpus sizes do not make a substantial contribution to the value of such metrics, when compared with smaller samples, and that adult and child corpuses have very similar profiles. Probabilities of occurrence for the phonemes, graphemes, and sonographs in this study may be downloaded from brm.psychonomic-journals.org/content/supplemental.
Matching stimuli across a range of influencing variables is no less important for studies of face recognition than it is for those of word processing. Whereas a number of corpora exist to allow experimenters to select a carefully controlled set of word stimuli, similar databases for famous faces do not exist. This article, therefore, provides researchers in the area of face recognition with a useful resource on which to base their stimulus selection. In the first phase of the investigation, British adults over 40 years of age were requested to generate the names of famous people (or celebrities) that they thought they would recognize and to write these down. The most frequently named celebrities were then rated by adults from the same age population for familiarity, distinctiveness, and age of acquisition. The result is a database of 696 famous people, with an indication of their relative eminence in the public consciousness and rated for these important variables. Phoneme counts are also provided for each famous person, together with family name frequency counts in the general population, where available. Materials and links may be accessed at www.psychonomic.org/archive.
A three-phased study was conducted in order to develop a standardized list of touch-related adjec-tives. The final list consisted of 306 words that were categorized in 440 instances according to the Le-derman and Klatzky (1987, 1990)dimensions of haptic properties (some words were classified in more than one dimension). The Kucera and Francis (1967)frequency of occurrence in written English for all words in the final list was also determined. A correlation was found between frequency of occurrence on the list and Kucera and Francis frequency. An analysis of the word dimensions and future applica-tions are discussed.
Planning, predicting, reasoning, and acting often depend crucially on the correct encoding and application of knowledge concerning the temporal and causal ordering of events. Yet no pictorial stimulus set is optimized for investigating the processing of temporal and causal order information. We introduce a novel stimulus set of 265 black-and-white line drawings depicting a diverse array of recognizable events. Most of the images in the stimulus set (N = 222) share a thematic or conceptual association with one other image in the set, and the stimuli were created and extensively normed such that the image pairs vary in the degrees to which they share a causal, ordered relation with one another. The stimuli were standardized in a series of normative tasks, including concept/noun/verb agreement, perceived frequency, visual similarity, and indexes of three features of causal associations between events (i.e., temporal proximity, exclusivity, and priority). Both younger adults (ages 18-30 years) and older adults (ages 60-80 years) contributed normative data, allowing for broad applications of the stimuli to the study of normal and age-related changes in the encoding, retention, and retrieval of information regarding temporal and causal order. Complete normative data sets are available in the online supplemental materials, and the full stimulus set is available by contacting the first author.
In this article, we describe the most extensive set of word associations collected to date. The database contains over 12,000 cue words for which more than 70,000 participants generated three responses in a multiple-response free association task. The goal of this study was (1) to create a semantic network that covers a large part of the human lexicon, (2) to investigate the implications of a multiple-response procedure by deriving a weighted directed network, and (3) to show how measures of centrality and relatedness derived from this network predict both lexical access in a lexical decision task and semantic relatedness in similarity judgment tasks. First, our results show that the multiple-response procedure results in a more heterogeneous set of responses, which lead to better predictions of lexical access and semantic relatedness than do single-response procedures. Second, the directed nature of the network leads to a decomposition of centrality that primarily depends on the number of incoming links or in-degree of each node, rather than its set size or number of outgoing links. Both studies indicate that adequate representation formats and sufficiently rich data derived from word associations represent a valuable type of information in both lexical and semantic processing.
Age of acquisition (AoA) is an important psycholinguistic variable that affects the speed and accuracy of lexical processing in tasks such as word naming, picture naming, and lexical decision. In the present work, we collected AoA ratings for 1,749 Portuguese words (nouns, verbs, adjectives, and adverbs), using a 9-point scale that was first proposed by Carroll and White (1973). We analyzed the relation between AoA ratings and other psycholinguistic variables (length measures, neighborhood density, written-word frequency, familiarity, imageability, and concreteness), and we assessed reliability by correlating our ratings with those from other databases presented for Portuguese, English, Spanish, and Italian. The full database can be downloaded from http://brm.psychonomic-journals.org/content/supplemental.
Two experiments attempted to resolve previous contradictory findings concerning developmental trends in false memories within the Deese-Roediger-McDermott (DRM) paradigm by using an improved methodology--constructing age-appropriate associative lists. The research also extended the DRM paradigm to preschoolers. Experiment 1 (N=320) included children in three age groups (preschoolers of 3-4 years, second-graders of 7-8 years, and preadolescents of 11-12 years) and adults, and Experiment 2 (N=64) examined preschoolers and preadolescents. Age-appropriate lists increased false recall. Although preschoolers had fewer false memories than the other age groups, they showed considerable levels of false recall when tested with age-appropriate materials. Results were discussed in terms of fuzzy-trace, source-monitoring, and activation frameworks.
Used correlation functions obtained in 2 experiments with undergraduate Os (N = 10) as a basis for describing human visual letter recognition. Visual images were filtered by means of autocorrelation for pattern information. This operation gave the relative visibilities or legibilities of the characters. The visual impressions were then cross-correlated with a set of memory records whose outputs described the relative probabilities that the stimulus was a given character. This operation described confusion errors. Finally "response bias" was described in terms of the reliability with which a memory record provides identification of a given stimulus. In these terms response bias represented an attempt by the recognition system to minimize errors in high-information responses, at the expense of producing more low-information responses as errors. (French summary)
This article introduces EsPal: a Web-accessible repository containing a comprehensive set of properties of Spanish words. EsPal is based on an extensible set of data sources, beginning with a 300 million token written database and a 460 million token subtitle database. Properties available include word frequency, orthographic structure and neighborhoods, phonological structure and neighborhoods, and subjective ratings such as imageability. Subword structure properties are also available in terms of bigrams and trigrams, biphones, and bisyllables. Lemma and part-of-speech information and their corresponding frequencies are also indexed. The website enables users either to upload a set of words to receive their properties or to receive a set of words matching constraints on the properties. The properties themselves are easily extensible and will be added over time as they become available. It is freely available from the following website: http://www.bcbl.eu/databases/espal/ .
Individual happiness is a fundamental societal metric. Normally measured through self-report, happiness has often been indirectly characterized and overshadowed by more readily quantifiable economic indicators such as gross domestic product. Here, we examine expressions made on the online, global microblog and social networking service Twitter, uncovering and explaining temporal variations in happiness and information levels over timescales ranging from hours to years. Our data set comprises over 46 billion words contained in nearly 4.6 billion expressions posted over a 33 month span by over 63 million unique users. In measuring happiness, we use a real-time, remote-sensing, non-invasive, text-based approach---a kind of hedonometer. In building our metric, made available with this paper, we conducted a survey to obtain happiness evaluations of over 10,000 individual words, representing a tenfold size improvement over similar existing word sets. Rather than being ad hoc, our word list is chosen solely by frequency of usage and we show how a highly robust metric can be constructed and defended.
This article describes a Windows program that enables users to obtain a broad range of statistics concerning the properties of word and nonword stimuli in Spanish, including word frequency, syllable frequency, bigram and biphone frequency, orthographic similarity, orthographic and phonological structure, concreteness, familiarity, imageability, valence, arousal, and age-of-acquisition measures. It is designed for use by researchers in psycholinguistics, particularly those concerned with recognition of isolated words. The program computes measures of orthographic similarity online, with respect to either a default vocabulary of 31,491 Spanish words or a vocabulary specified by the user. In addition to providing standard orthographic and phonological neighborhood measures, the program can be used to obtain information about other forms of orthographic similarity, such as transposed-letter similarity and embedded-word similarity. It is available, free of charge, from the following Web site: www.maccs.mq.edu.au/-colin/B-Pal.
In this article, normative data on the familiarity and difficulty of 196 single-solution Spanish word fragments are presented. The database includes the following indices: difficulty, familiarity, frequency, number of meanings, number of letters given in the fragment, first and/or last letters given, and ratio of letters to blanks. A factor analysis was performed on difficulty, and two factors were obtained. Frequency, familiarity, and number of meanings loaded highly on the first factor, which we consider to measure lexical processes, whereas number of letters in the fragment, first and/or last letters given, and ratio of letters to blanks loaded highly on the second factor, which we judge to be determined by perceptual information. Regression analyses using factor scores as predictors showed that both factors accounted for a significant part of the completion probability scores. The full set of these norms may be downloaded from the Psychonomic Society Web archive at
Ratings of concreteness and picturability and production data for meaningfulness of 310 words were gathered from 207 6th-grade children and 265 college adults. Adults also provided ratings of imagery value. Correlations among the various stimulus attributes indicated that for adults, imagery, concreteness, and picturability were overlapping attributes. For children, however, the attributes of concreteness and picturability did not overlap as much.
This paper describes a Japanese logographic character (kanji) frequency list, which is based on an analysis of the largest recently available corpus of Japanese words and characters. This corpus comprised a full year of morning and evening editions of a major newspaper, containing more than 23 million kanji characters and more than 4,000 different kanji characters. This paper lists the 3,000 most frequent kanji characters, as well as an analysis of kanji usage and correlations between the present list and previous Japanese frequency lists. The authors believe that the present list will help researchers more accurately and efficiently control the selection of kanji characters in cognitive science research and interpret related psycholinguistic data.
A set of semantically neutral sentences and derived pseudosentences was produced by two native European Portuguese speakers varying emotional prosody in order to portray anger, disgust, fear, happiness, sadness, surprise, and neutrality. Accuracy rates and reaction times in a forced-choice identification of these emotions as well as intensity judgments were collected from 80 participants, and a database was constructed with the utterances reaching satisfactory accuracy (190 sentences and 178 pseudosentences). High accuracy (mean correct of 75{\%} for sentences and 71{\%} for pseudosentences), rapid recognition, and high-intensity judgments were obtained for all the portrayed emotional qualities. Sentences and pseudosentences elicited similar accuracy and intensity rates, but participants responded to pseudosentences faster than they did to sentences. This database is a useful tool for research on emotional prosody, including cross-language studies and studies involving Portuguese-speaking participants, and it may be useful for clinical purposes in the assessment of brain-damaged patients. The database is available for download from http://brm.psychonomic-journals.org/content/supplemental.
This study presents Portuguese category norms for children of three different age groups: preschoolers (3- to 4-year-olds), second graders (7- to 8-year-olds), and preadolescents (11- to 12-year-olds). Three hundred Portuguese children (100 in each group) completed an exemplar-generation task. Preschoolers generated exemplars for 13 categories, second graders generated exemplars for 17 categories, and preadolescents generated exemplars for 21 categories. For each group, responses within each category were organized according to frequency of production in order to derive exemplar-production norms for sets of tested categories. The results also included information about the number of responses and exemplars, idiosyncratic and inappropriate responses, and commonality and diversity indexes for all the categories. A comparison of these children's norms with the Portuguese adult norms was also presented. The full set of norms may be downloaded from www.psychonomic.org/archive.
Age of acquisition is one of the most important variables in picture naming. For this reason, a large number of findings concerning age-of-acquisition data have been published in recent years in a number of different languages. In this article, objective age-of-acquisition data in Spanish for 328 pictures were collected from a pool of 760 children, half of whom were boys and the other half girls. A total of 246 pictures were selected from the Snodgrass and Vanderwart (1980) set, and 82 were new pictures. Like the results of other studies, we found that objective age of acquisition correlates less than rated age of acquisition with familiarity and frequency, which indicates that the objective measure is less contaminated by other variables than are rated estimates. A very high correlation was obtained between the norms from this study and those published in English, French, Icelandic, and Italian. These norms will be very useful to Spanish psycholinguists and clinicians. Related materials may be downloaded from the Psychonomic Society Web archive at www.psychonomic.org/archive/.
Two-word familiarity sets were measured in different years (1995 and 2002) and places (Kanto and Kinki, in Japan) for a large number of Japanese words, to examine the reliability of familiarity ratings. The correlation between the word familiarities of the two sets was extremely high (r ? .958, N ? 10,515). It is suggested that familiarity rating, at least for ordinary words found in a dictionary, is very reliable and not greatly affected by differences in years and places.
Collected normative data for 254 line drawings from the set used by J. G. Snodgrass and M. Vanderwart (see record 1981-06756-001) to be used in research with Spanish-speaking samples. 261 Spanish-speaking Ss participated in 1 of 6 tasks: name agreement, familiarity, complexity, image agreement, picture-name agreement, and image variability. Each S responded to every drawing. Results are compared to those obtained by Snodgrass and Vanderwart from English-speaking Ss. There were small but significant differences for familiarity and complexity. The English-speaking sample rated the pictures as more familiar; the Spanish Ss judged the pictures as slightly more simple. The evidence justifies the statement that normative data of cognitive stimuli cannot be taken into another language directly, because object names common in one language may not be so in another, or objects that have a specific name in one language may have a generic name in another. (PsycINFO Database Record (c) 2007 APA, all rights reserved)
The LEXIN database offers psycholinguistic indexes of the 13,184 different words (types) computed from 178,839 occurrences of these words (tokens) contained in a corpus of 134 beginning readers widely used in Spain. This database provides four statistical indicators: F (overall word frequency), D (index of dispersion across selected readers), U (estimated frequency per million words), and SFI (standard frequency index). It also gives information about the number of letters, syntactic category, and syllabic structure of the words included. To facilitate comparisons, LEXIN provides data from LEXESP's (Sebasti{\'{a}}n-Gall{\'{e}}s, Mart{\'{i}}, Cuetos, {\&} Carreiras, 2000), Alameda and Cuetos's (1995), and Mart{\'{i}}nez and Garc{\'{i}}a's (2004) Spanish adult psycholinguistic frequency databases. Access to the LEXIN database is facilitated by a computer program. The LEXIN program allows for the creation of word lists by letting the user specify searching criteria. LEXIN can be useful for researchers in cognitive psychology, particularly in the areas of psycholinguistics and education.
Two databases of Spanish surface word forms are presented. Surface word forms are words considered as orthographically or phonologically specified without reference to their meaning or syntactic category. The databases are based on the productive written vocabulary of children between the ages of 6 and 10 years. Statistical and structural information is presented concerning surface word-form frequency, consonant-vowel (CV) structure, number of syllables, syllables, syllable CV structure, and subsyllabic units. LEX I was intended to aid in the study of reading processes. Entries were orthographic surface word forms; words were divided in their components following orthographic criteria. LEX II was designed for spoken language research. Accordingly, words were transcribed phonologically and phonological criteria were applied in extracting the internal units. Information about stress location was also provided. Together, LEX I and LEX II represent a useful tool for psycholinguists interested in the study of people acquiring Spanish as a first or foreign language and of Spanish-speaking populations in general
Frequency of occurrence is an important attribute of lexical units, and one that is widely used in psychological research and theorization. Although printed frequency norms have long been available for Spanish, and subtitle-based norms have more recently been published, oral frequency norms have not been systematically compiled for a representative set of words. In this study, a corpus of over three million units, representing present-day use of the language in Spain, was used to derive a frequency count of spoken words. The corpus consisted of 913 separate documents that contained transcriptions of oral recordings obtained in a wide variety of situations, mostly radio and television programs. The resulting database, containing absolute and relative frequency values for 67,979 orally produced words, is presented. Validity analyses showed significant correlations of oral frequency with other frequency measures and suggest that oral frequency can predict some types of lexical processing with the same or higher levels of precision, when contrasted with text- or subtitle-based frequencies. In conclusion, we discuss ways in which these oral frequency norms can be put to use. The norms can be downloaded from www.springerlink.com.
A normative study was conducted using the Deese/Roediger-McDermott paradigm (DRM) to obtain false recognition for 60 six-word lists in Spanish, designed with a completely new methodology. For the first time, lists included words (e.g., bridal, newlyweds, bond, commitment, couple, to marry) simultaneously associated with three critical words (e.g., love, wedding, marriage). Backward associative strength between lists and critical words was taken into account when creating the lists. The results showed that all lists produced false recognition. Moreover, some lists had a high false recognition rate (e.g., 65{\%}; jail, inmate, prison: bars, prisoner, cell, offender, penitentiary, imprisonment). This is an aspect of special interest for those DRM experiments that, for example, record brain electrical activity. This type of list will enable researchers to raise the signal-to-noise ratio in false recognition event-related potential studies as they increase the number of critical trials per list, and it will be especially useful for the design of future research.