1396 norm sets
This article presents the Spanish adaptation of the Affective Norms for English Words (ANEW; Bradley {\&} Lang, 1999). The norms are based on 720 participants' assessments of the translation into Spanish of the 1,034 words included in the ANEW. The evaluations were done in the dimensions of valence, arousal and dominance using the Self-Assessment Manikin (SAM). Apart from these dimensions, five objective (number of letters, number of syllables, grammatical class, frequency and number of orthographic neighbors) and three subjective (familiarity, concreteness and imageability) psycholinguistic indexes are included. The Spanish adaptation of ANEW can be downloaded at www.psychonomic.org.
The purpose of the present investigation was to replicate and extend the International Affective Picture System norms (Ito, Cacioppo, {\{}{\&}{\}} Lang, 1998; Lang, Bradley, {\{}{\&}{\}} Cuthbert, 1999). These norms were developed to provide researchers with photographic slides that varied in emotional evocation, especially arousal and valence. In addition to collecting rating data on the dimensions of arousal and valence, we collected data on the dimensions of consequentiality, meaningfulness, familiarity, distinctiveness, and memorability. Furthermore, we collected ratings on the primary emotions of happiness, surprise, sadness, anger, disgust, and fear. A total of 1,302 participants were tested in small groups. The participants in each group rated a subset of 18 slides on 14 dimensions. Ratings were obtained on 703 slides. The means and standard deviations for all of the ratings are provided. We found our valence ratings to be similar to the previous norms. In contrast, our participants were more likely to rate the slides as less arousing than in the previous norms. The mean ratings on the remaining 12 dimensions were all below the midpoint of the 9-point Likert scale. However, sufficient variability in ratings across the slides indicates that selecting slides on the basis of these variables is feasible. Overall, the present ratings should allow investigators to use these norms for research purposes, especially in research dealing with the interrelationships among emotion and cognition. The means and standard deviations for emotions may be downloaded as an Excel spreadsheet from www.psychonomic.org/archive.
The memory block effect (MBE) occurs when orthographically similar words inhibitretrieval. Previous studies have published 55 different stimuli that produce the MBE in word fragment completion. This small number of stimuli constrains experimental designs, presents serious obstacles for using neuroimaging to elucidate neural substrates of blocking, and raises concern that the MBE is limited to a particular group of words. A pool of 315 stimulus words was tested in a traditional MBE paradigm, and the results demonstrated that the MBE generalizes to other stimuli. This study also expands the number of stimuli that produce the MBE because 185 new stimuli produced blocking effects. As a result, the current list of 240 MBE stimuli can be used for word fragment research including cognitive neuroscience investigations of retrieval inhibition. A table of MBE stimuli is available in an archived appendix that can be downloaded from www.psychonomic.org/archive.
The majority of research on the acquisition of spoken language has focused on language production, due to difficulties in the assessment of comprehension. A primary limitation to comprehension assessment is maintaining the interest and attention of younger infants. We have developed an assessment procedure that addresses the need for an extensive performance-based measure of comprehension in the 2nd year of life. In the interest of developing an engaging approach that takes into account infants' limited attention capabilities, we designed an assessment based on touchscreen technology. This approach builds upon prior research by combining standardization and complexity with an engaging infant-friendly interface. Data suggest that the touchscreen procedure is effective in eliciting and maintaining infant attention and will yield more extensive and reliable estimates of early comprehension than do other procedures. The software to implement the assessment is available free of charge for academic purposes.
Much of the power of neural network modeling for language use and acquisition derives from a reliance on statistical regularities implicit in the phonological properties of words. Researchers have devised several methods for representing the phonology of words, but these methods are often either unable to represent realistically sized lexicons or inadequate in the ways they represent individual words. In this paper, we present a new phonological pattern generator (PatPho) that allows connectionist modelers to derive accurate phonological representations of the English lexicon. PatPho not only generates phonological patterns that can scale up to realistically sized lexicons, but also accurately and parsimoniously captures the similarity structures of the phonology of monosyllabic and multisyllabic words.
We present modality exclusivity norms for 400 randomly selected noun concepts, for which participants provided perceptual strength ratings across five sensory modalities (i.e., hearing, taste, touch, smell, and vision). A comparison with previous norms showed that noun concepts are more multimodal than adjective concepts, as nouns tend to subsume multiple adjectival property concepts (e.g., perceptual experience of the concept baby involves auditory, haptic, olfactory, and visual properties, and hence leads to multimodal perceptual strength). To show the value of these norms, we then used them to test a prediction of the sound symbolism hypothesis: Analysis revealed a systematic relationship between strength of perceptual experience in the referent concept and surface word form, such that distinctive perceptual experience tends to attract distinctive lexical labels. In other words, modality-specific norms of perceptual strength are useful for exploring not just the nature of grounded concepts, but also the nature of form-meaning relationships. These norms will be of benefit to those interested in the representational nature of concepts, the roles of perceptual information in word processing and in grounded cognition more generally, and the relationship between form and meaning in language development and evolution.
Color is undeniably important to object representations, but so too is the ability of context to alter the color of an object. The present study examined how implied perceptual information about typical and atypical colors is represented during language comprehension. Participants read sentences that implied a (typical or atypical) color for a target object and then performed a modified Stroop task in which they named the ink color of the target word (typical, atypical, or unrelated). Results showed that color naming was facilitated both when ink color was typical for that object (e.g., bear in brown ink) and when it matched the color implied by the previous sentence (e.g., bear in white ink following Joe was excited to see a bear at the North Pole). These findings suggest that unusual contexts cause people to represent in parallel both typical and scenario-specific perceptual information, and these types of information are discussed in relation to the specialization of perceptual simulations.
Cognitive models of language processing in English are founded on norms for word properties, but their universality is now being explored across different writing scripts and subject groups. Although Chinese characters are popular for this comparative work, their salient properties remain ill defined or poorly controlled. We describe how norms for semantic and phonetic regularity in Mandarin can be calibrated on a regional basis. The rating data that we present from China, Singapore, and Taiwan also illustrate why the diversity of both oral and written forms of Chinese should be considered in future empirical work.
The Character-Component Analysis Toolkit (C-CAT) software was designed to assist researchers in constructing experimental materials using traditional Chinese characters. The software package contains two sets of character stocks: one suitable for research using literate adults as subjects and one suitable for research using schoolchildren as subjects. The software can identify linguistic properties, such as the number of strokes contained, the character-component pronunciation regularity, and the arrangement of character components within a character. Moreover, it can compute a character's linguistic frequency, neighborhood size, and phonetic validity with respect to a user-selected character stock. It can also search the selected character stock for similar characters or for character components with user-specified linguistic properties.
The use of multilevel modeling is presented as an alternative to separate item and subject ANOVAs (F1 x F2) in psycholinguistic research. Multilevel modeling is commonly utilized to model variability arising from the nesting of lower level observations within higher level units (e.g., students within schools, repeated measures within individuals). However, multilevel models can also be used when two random factors are crossed at the same level, rather than nested. The current work illustrates the use of the multilevel model for crossed random effects within the context of a psycholinguistic experimental study, in which both subjects and items are modeled as random effects within the same analysis, thus avoiding some of the problems plaguing current approaches.
In this article, we present a new lexical database for Modern Standard Arabic: Aralex. Based on a contemporary text corpus of 40 million words, Aralex provides information about (1) the token frequencies of roots and word patterns, (2) the type frequency, or family size, of roots and word patterns, and (3) the frequency of bigrams, trigrams in orthographic forms, roots, and word patterns. Aralex will be a useful tool for studying the cognitive processing of Arabic through the selection of stimuli on the basis of precise frequency counts. Researchers can use it as a source of information on natural language processing, and it may serve an educational purpose by providing basic vocabulary lists. Aralex is distributed under a GNU-like license, allowing people to interrogate it freely online or to download it from www.mrc-cbu.cam.ac.uk:8081/aralex.online/login.jsp.
This article present the Spanish assessments of the 111 sounds included in the International Affective Digitized Sounds (IADS; Bradley {\&} Lang, 1999b). The sounds were evaluated by 159 participants in the dimensions of valence, arousal, and dominance, using a computer version of the Self-Assessment Manikin (Bradley {\&} Lang, 1994). Results are compared with those obtained in the American version of the IADS, as well as in the Spanish adaptations of the International Affective Picture System (P. J. Lang, Bradley, {\&} Cuthbert, 1999; Molt{\'{o}} et al., 1999) and the Affective Norms for English Words (Bradley {\&} Lang, 1999a; Redondo, Fraga, Padr{\'{o}}n, {\&} Comesa{\~{n}}a, 2007).
The CFVlexvar.xls database includes imageability, frequency, and grammatical properties of the first words acquired by Italian children. For each of 519 words that are known by children 18-30 months of age (taken from Caselli {\&} Casadio's, 1995, Italian version of the MacArthur Communicative Development Inventory), new values of imageability are provided and values for age of acquisition, child written frequency, and adult written and spoken frequency are included. In this article, correlations among the variables are discussed and the words are grouped into grammatical categories. The results show that words acquired early have imageable referents, are frequently used in the texts read and written by elementary school children, and are frequent in adult written and spoken language. Nouns are acquired earlier and are more imageable than both verbs and adjectives. The composition in grammatical categories of the child's first vocabulary reflects the composition of adult vocabulary. The full set of these norms can be downloaded from www.psychonomic.org/archive/
In this study, we provide normative data for objects in a set of 80 digital color pictures (e.g., nature scenes, human activities, cartoon characters, magazine covers). In Experiment 1, four objects in each picture were rated by 48 observers on a 6-point Likert scale for their relevance to the overall meaning of the scene. In Experiment 2, Salience Toolbox software (Walther {\&} Koch, 2006) provided additional information about whether the four relevance-rated objects were located in areas that were high or low in visual salience. Brief descriptions of the four objects, their locations in the picture, their categorizations as high or low in salience, the means and standard deviations of their relevance ratings, and statistical analyses specifying which pairs of objects in a picture differed significantly on their relevance to the meaning of the scene are given in the Appendix. An example is provided of how the pictures could be used to create stimuli for a change blindness task in which detection of item onset versus offset is contrasted for low-relevance and high-relevance features. The 80 pictures are accessible as jpg files from the first author's Web site at http://marcellm.people.cofc.edu/research.htm.
In this article, we present a database of orthographic neighbors for words that Spanish children read during elementary education. The reference dictionary for lexical entries and frequencies (which had its origin in Mart{\'{i}}nez {\&} Garc{\'{i}}a, 2004) comprises approximately 100,000 words and is the result of accumulating the words read by a sample of children from first to sixth grades. Using the criterion for orthographic neighbors described by Coltheart, Davelaar, Jonasson, and Besner (1977), we present basic statistics related to neighborhood size as a function of the positions of divergent letters, the cumulative frequency of the neighbors, and the numbers of neighbors of higher, lower, and equal frequency. We also attempt to illustrate and unravel the nature of the relationships among the variables neighborhood size, length, and frequency in the distribution of neighbors. The database described in this article is available at www.psychonomic.org/archive.
The main objective of this study is to report rated age of acquisition (AoA) norms for 834 nouns in Portuguese (European). AoA ratings were collected on a 7 point scale, generally following Gilhooly and Logie (1980) procedure with an 8 extra point for "don't know the word" answers. Results were analyzed considering AoA ratings and their standard deviations and considering the relationship between AoA ratings and other psycholinguistic variables (imageability, familiarity, written word frequency, concreteness, number of syllables and number of words). AoA ratings and their standard deviations were significantly and positively correlated, with early acquired word ratings showing higher agreement. Correlation and multiple regression analyses confirmed the major contribution of imageability and familiarity to AoA ratings obtained in other languages. The full database of AoA ratings and other psycholinguistic variables may be downloaded from www.psychonomic.org/archive or www.fpce.ul.pt/pessoal/ulfpfred/aoa.htm.
The study of the cognitive processes in the production of language demands careful selection of stimuli and requires normative databases. The main goal of the present research was to collect normative data for the set of 400 figures taken from Cycowicz, Friedman, Rothstein, and Snodgrass (1997; including the 260 figures of Snodgrass {\&} Vanderwart, 1980) using a sample of native Argentinean Spanish speakers. The pictures have been standardized on the following variables: name agreement, image agreement, familiarity, visual complexity, image variability, age of acquisition, and word association. The obtained norms were compared with the normative data of other studies in Spanish, English, and French. This comparison highlights the variability of some of the measures (e.g., name agreement in naming and verbal association) across the different studies and confirms the necessity of elaborating specific norms that are adapted to the studied population's linguistic and sociocultural context. The norms described may be downloaded as supplemental materials for this article from http://brm.psychonomic-journals.org/content/supplemental.
Sound events are sequences of closely grouped and temporally related environmental sounds that tell a story or establish a sense of place. The goal of our project was to create a set of sound events depicting various scenarios (such as a car accident, cooking breakfast, and walking outdoors) and to gather normative data about how people understand them. Samples of college students listened to 22 sound events over headphones in three self-paced, computer-based studies. In the Identification Task, 43 participants used text boxes to type descriptions of what was happening in the sound events. In the Rating Task, 39 participants used Likert scales to rate the sound events on the attributes of familiarity, complexity, and pleasantness. In the Memory Task, 42 participants answered two multiple-choice questions immediately after listening to each sound event. Detailed tables are provided for the following: (1) Description of the sound events and their components; (2) accuracy and response time measurements for each of the 22 sound events across the three studies; and (3) rank-orderings of the sound events by ease of identification, recognition of details, and rated familiarity, complexity, and pleasantness. Digital files of the stimuli, which may be of interest to auditory cognition researchers and clinical neuropsychologists, may be downloaded from either www.psychonomic.org/archive or www.cofc.edu/-marcellm/sound event studies/sndevent.htm.
Communication using icons is now commonplace. It is therefore important to understand the processes involved in icon comprehension and the stimulus cues that individuals utilize to facilitate identification. In this study, we examined predictors of icon identification as participants gained experience with icons over a series of learning trials. A dynamic pattern of findings emerged in which the primary predictors of identification changed as learning progressed. In early learning trials, semantic distance (the closeness of the relationship between icon and function) was the best predictor of performance, accounting for up to 55{\%} of the variance observed, whereas familiarity with the function was more important in later trials. Other stimulus characteristics, such as our familiarity with the graphic in the icon and its concreteness, were also found to be important for icon design. The theoretical implications of these findings are discussed, with particular emphasis on the parallels with picture naming. The icon identification norms from this study may be downloaded from brm.psychonomic-journals.org/content/supplemental.
The strength-sampling model of free association (Nelson, McEvoy, {\&} Dennis, 2000) claims that the probability of word association in free-association norms results from a sampling process. For a given cue word, each response word has an underlying distribution of strength values. In the free-association task, presentation of the cue word activates a random sample of strengths, one for each response. The highest strength wins, and its response is reported. In the present work, gradient descent was used to compute the theoretical mean strengths for each cue-response pair in the Nelson, McEvoy, and Schreiber (2004) norms. The resulting database may be downloaded from www.psychonomic.org/archive/.
Semantic features have provided insight into numerous behavioral phenomena concerning concepts, categorization, and semantic memory in adults, children, and neuropsychological populations. Numerous theories and models in these areas are based on representations and computations involving semantic features. Consequently, empirically derived semantic feature production norms have played, and continue to play, a highly useful role in these domains. This article describes a set of feature norms collected from approximately 725 participants for 541 living (dog) and nonliving (chair) basic-level concepts, the largest such set of norms developed to date. This article describes the norms and numerous statistics associated with them. Our aim is to make these norms available to facilitate other research, while obviating the need to repeat the labor-intensive methods involved in collecting and analyzing such norms. The full set of norms may be downloaded from www.psychonomic.org/archive.
Two sentences are paraphrases if their meanings are equivalent but their words and syntax are different. Paraphrasing can be used to aid comprehension, stimulate prior knowledge, and assist in writing-skills development. As such, paraphrasing is a feature of fields as diverse as discourse psychology, composition, and computer science. Although automated paraphrase assessment is both commonplace and useful, research has centered solely on artificial, edited paraphrases and has used only binary dimensions (i.e., is or is not a paraphrase). In this study, we use an extensive database (N = 1,998) of natural paraphrases generated by high school students that have been assessed along 10 dimensions (e.g., semantic completeness, lexical similarity, syntactical similarity). This study investigates the components of paraphrase quality emerging from these dimensions and examines whether computational approaches can simulate those human evaluations. The results suggest that semantic and syntactic evaluations are the primary components of paraphrase quality, and that computationally light systems such as latent semantic analysis (semantics) and minimal edit distances (syntax) present promising approaches to simulating human evaluations of paraphrases. (PsycINFO Database Record (c) 2012 APA, all rights reserved). (journal abstract)
Two studies were conducted in which human participants rated pairs of words according to the perceived degree to which the words' referents shared semantic features. The participants found the task intuitive, simple, and quick to complete. The ratings were reliable and valid. Interrater and interstudy correlations were high, and ratings were good predictors of known feature overlap values obtained from existing semantic feature norms. (PsycINFO Database Record (c) 2006 APA ) (journal abstract)
Ratings of realism, masculinity, race, and racial stereotypy were collected on a set of computer-generated faces representing European, South East Asian, and African American ethnicities. To determine if these faces are processed in the same way as photographs of real faces, we demonstrated with these faces superior memory performance for upright faces over inverted faces (the face inversion effect). Further, in observers of European decent, we found both superior memory for European faces and a larger inversion effect for European than African American faces. Based on these results, we believe that this set of faces may be of use in perceptual investigations in which race is a critical manipulation.
Four experiments were conducted to assess two models of topic sentencehood identification: the derived model and the free model. According to the derived model, topic sentences are identified in the context of the paragraph and in terms of how well each sentence in the paragraph captures the paragraph's theme. In contrast, according to the free model, topic sentences can be identified on the basis of sentential features without reference to other sentences in the paragraph (i.e., without context). The results of the experiments suggest that human raters can identify topic sentences both with and without the context of the other sentences in the paragraph. Another goal of this study was to develop computational measures that approximated each of these models. When computational versions were assessed, the results for the free model were promising; however, the derived model results were poor. These results collectively imply that humans' identification of topic sentences in context may rely more heavily on sentential features than on the relationships between sentences in a paragraph.
There is increasing interest in the role that manipulability plays in processing objects. To date, Magni{\'{e}}, Besson, Poncet, and Dolisi's (2003) manipulability ratings, based on the degree to which objects can be uniquely pantomimed, have been the reference point for many studies. However, these ratings do not fully capture some relevant dimensions of manipulability, including whether an object is graspable and the extent to which functional motor associations above and beyond graspability are present. To address this, we collected ratings of these dimensions, in addition to ratings of familiarity and age of acquisition (AoA), for a set of 320 black-and-white photographs of objects. Familiarity and AoA ratings were highly correlated with previously reported ratings of the same dimensions (r = .853, p {\textless} .001, and r = .771, p {\textless} .001, respectively), validating the present norms. Grasping and functional use ratings, in contrast, were more moderately correlated with Magni{\'{e}} et al.'s pantomime manipulability ratings (r = .507, p {\textless} .001). These results were taken as evidence that the new manipulability ratings collected in this research capture distinct aspects of object manipulability. The complete stimuli and norms from this study may be downloaded from http://brm.psychonomic-journals.org/content/supplemental.
A data set is described that includes eight variables gathered for 13 common superordinate natural language categories and a representative set of 338 exemplars in Dutch. The category set contains 6 animal categories (reptiles, amphibians, mammals, birds, fish, and insects), 3 artifact categories (musical instruments, tools, and vehicles), 2 borderline artifact-natural-kind categories (vegetables and fruit), and 2 activity categories (sports and professions). In an exemplar and a feature generation task for the category nouns, frequency data were collected. For each of the 13 categories, a representative sample of 5-30 exemplars was selected. For all exemplars, feature generation frequencies, typicality ratings, pairwise similarity ratings, age-of-acquisition ratings, word frequencies, and word associations were gathered. Reliability estimates and some additional measures are presented. The full set of these norms is available in Excel format at the Psychonomic Society Web archive, www.psychonomic.org/archive/.
Sentence completion norms are a valuable resource for researchers interested in studying the effects of context on word recognition processes. Norms for 112 Spanish sentences were compiled with the use of experimental software accessed over the World-Wide Web. Several measures summarizing the distribution of responses for each sentence are reported, including Schwanenflugel's (1986) multiple-production measure of sentence constraint strength, the type-token ratio, and the information-theoretic measure of redundancy. The complete set of completion norms is available at http://www.ling.ed.ac.uk/{\~{}}monica/spanish{\_}completion{\_}norms.html.
We provide imageability estimates for 3,000 disyllabic words (as supplementary materials that may be downloaded with the article from www.springerlink.com ). Imageability is a widely studied lexical variable believed to influence semantic and memory processes (see, e.g., Paivio, 1971). In addition, imageability influences basic word recognition processes (Plaut, McClelland, Seidenberg, {\&} Patterson, 1996). In fact, neuroimaging studies have suggested that reading high- and low-imageable words elicits distinct neural activation patterns for the two types e.g., Bedny {\&} Thompson-Schill (Brain and Language 98:127-139, 2006; Graves, Binder, Desai, Conant, {\&} Seidenberg NeuroImage 53:638-646, 2010). Despite the usefulness of this variable, imageability estimates have not been available for large sets of words. Furthermore, recent megastudies of word processing e.g., Balota et al. (Behavior Research Methods 39:445-459, 2007) have expanded the number of words that interested researchers can select according to other lexical characteristics (e.g., average naming latencies, lexical decision times, etc.). However, the dearth of imageability estimates (as well as those of other lexical characteristics) limits the items that researchers can include in their experiments. Thus, these imageability estimates for disyllabic words expand the number of words available for investigations of word processing, which should be useful for researchers interested in the influences of imageability both as an input and as an outcome variable.
Syllogistic reasoning, in which people identify conclusions from quantified premise pairs, remains a benchmark task whose patterns of data must be accounted for by general theories of deductive reasoning. However, psychologists have confined themselves to administering only the 64 premise pairs historically identified by Aristotle. By utilizing all combinations of negations, the present article identifies an expanded set of 576 premise pairs and gives the valid conclusions that they support. Many of these have interesting properties, and the identification of predictions and their verification will be an important next step for all proponents of such theories.
This study examines the relationship between the linguistic characteristics of body paragraphs of student essays and the total number of paragraphs in the essays. Results indicate a significant relationship between the total number of paragraphs and a variety of linguistic characteristics known to affect student essay scores. These linguistic characteristics (e.g., semantic overlap, syntactic complexity) contribute to two underlying factors (i.e., textual cohesion and difficulty) that are used as dependent variables in mixed-effect models. Results suggest that student essays with 5-8 paragraphs tend to be more linguistically consistent than student essays with 3, 4, and 9 paragraphs. Essays with totals of 5-8 paragraphs, considered by many educators to contain an optimal number of paragraphs, may include functionally and structurally similar paragraphs. These findings could aid writing researchers and educators in obtaining a clearer view of the relationship between the total number of paragraphs comprising an essay and the linguistic characteristics that affect essay evaluation. Consequently, writing interventions may become better equipped to pinpoint student difficulties and facilitate student writing skills by providing more detailed and informed feedback.
Although many individual speech contrasts pairs have been studied within the cross-language literature, no one has created a comprehensive and systematic set of such stimuli. This article justifies and details an extensive set of contrast pairs for Mandarin Chinese and American English. The stimuli consist of 180 pairs of CVC syllables recorded in two tokens each (720 syllables total). Between each CVC pair, two of the segments are identical, whereas the third differs in that a segment drawn from a "native" phonetic category (either Mandarin, English, or both) is partnered with a segment drawn from a "foreign" phonetic category (nonnative to Mandarin, English, or both). Each contrast pair differs by a minimal phonetic amount and constitutes a meaningful contrast among the world's languages (as cataloged in the UCLA Phonological Segment Inventory Database of 451 languages). The entire collection of phonetic differences envelops Mandarin and English phonetic spaces and generates a range of phonetic discriminability. Contrastive segments are balanced through all possible syllable positions, with noncontrastive segments being filled in with other "foreign" segments. Although intended to measure phonetic perceptual sensitivity among adult speakers of the two languages, these stimuli are offered here to all for similar or for altogether unrelated investigations.
A strong body of work has explored the interaction between visual perception and language comprehension; for example, recent studies exploring predictions from embodied cognition have focused particularly on the common representation of sensory-motor and semantic information. Motivated by this background, we provide a set of norms for the axis and direction of motion implied in 299 English verbs, collected from approximately 100 native speakers of British English. Until now, there have been no freely available norms of this kind for a large set of verbs that can be used in any area of language research investigating the semantic representation of motion. We have used these norms to investigate the interaction between language comprehension and low-level visual processes involved in motion perception, validating the norming procedure's ability to capture the motion content of individual verbs. Supplemental materials for this study may be downloaded from brm.psychonomic-journals.org/content/supplemental.
Faces constitute a unique and widely used category of stimuli. In spite of their importance, there are few collections of faces for use in research, none of which adequately represent the different ages of faces across the lifespan. This lack of a range of ages has limited the majority of researchers to using predominantly young faces as stimuli even when their hypotheses concern both young and old participants. We describe a database of 575 individual faces ranging from ages 18 to 93. Our database was developed to be more representative of age groups across the lifespan, with a special emphasis on recruiting older adults. The resulting database has faces of 218 adults age 18-29, 76 adults age 30-49, 123 adults age 50-69, and 158 adults age 70 and older. These faces may be acquired for research purposes from http://agingmind.cns.uiuc.edu/facedb/. This will allow researchers interested in using facial stimuli access to a wider age range of adult faces than has previously been available.
This article presents norms of valence/pleasantness, activity/arousal, power/dominance, and age of acquisition for 4,300 Dutch words, mainly nouns, adjectives, adverbs, and verbs. The norms are based on ratings with a 7-point Likert scale by independent groups of students from two Belgian (Ghent and Leuven) and two Dutch (Rotterdam and Leiden-Amsterdam) samples. For each variable, we obtained high split-half reliabilities within each sample and high correlations between samples. In addition, the valence ratings of a previous, more limited study (Hermans {\&} De Houwer, Psychologica Belgica, 34:115-139, 1994) correlated highly with those of the present study. Therefore, the new norms are a valuable source of information for affective research in the Dutch language.
In this article, we present a new lexical database for French: Lexique. In addition to classical word information such as gender, number, and grammatical category, Lexique includes a series of interesting new characteristics. First, word frequencies are based on two cues: a contemporary corpus of texts and the number of Web pages containing the word. Second, the database is split into a graphemic table with all the relevant frequencies, a table structured around lemmas (particularly interesting for the study of the inflectional family), and a table about surface frequency cues. Third, Lexique is distributed under a GNU-like license, allowing people to contribute to it. Finally, a metasearch engine, Open Lexique, has been developed so that new databases can be added very easily to the existing ones. Lexique can either be downloaded or interrogated freely from http://www.lexique.org.
Word stem completion tasks involve showing participants a number of words and then later asking them to complete word stems to make a full word. If the stem is completed with one of the studied words, it indicates memory. It is a test widely used to assess both implicit and explicit forms of memory. An important aspect of stimulus selection is that target words should not frequently be generated spontaneously from the word stem, to ensure that production of the word really represents memory. In this article, we present a database of spontaneous stem completion rates for 395 stems from a group of 80 British undergraduate psychology students. It includes information on other characteristics of the words (word frequency, concreteness, imageability, age of acquisition, common part of speech, and number of letters) and, as such, can be used to select suitable words to include in a stem completion task. Supplemental materials for this article may be downloaded from http://brm.psychonomic-journals.org/content/supplemental.
Hyperspace analog to language (HAL) is a high-dimensional model of semantic space that uses the global co-occurrence frequency of words in a large corpus of text as the basis for a representation of semantic memory. In the original HAL model, many parameters were set without any a priori rationale. We have created and publicly released a computer application, the High Dimensional Explorer (HiDEx), that makes it possible to systematically alter the values of these parameters to examine their effect on the co-occurrence matrix that instantiates the model. We took an empirical approach to understanding the influence of the parameters on the measures produced by the models, looking at how well matrices derived with different parameters could predict human reaction times in lexical decision and semantic decision tasks. New parameter sets give us measures of semantic density that improve the model's ability to predict behavioral measures. Implications for such models are discussed.
This study provides Japanese normative measures for 359 line drawings, including 260 pictures (44 redrawn) taken from Snodgrass and Vanderwart (1980). The pictures have been standardized on voice key naming times, name agreement, age of acquisition, and familiarity. The data were compared with American, Spanish, French, and Icelandic samples reported in previous studies. In general, the correlations between variables in the present study and those in the other studies were relatively high, except for name agreement. Naming times were predicted in multiple regression analyses by name agreement. The full set of the norms and the new pictures may be downloaded from www.psychonomic.org/archive/.
Preexisting word knowledge is accessed in many cognitive tasks, and this article offers a means for indexing this knowledge so that it can be manipulated or controlled. We offer free association data for 72,000 word pairs, along with over a million entries of related data, such as forward and backward strength, number of competing associates, and printed frequency. A separate file contains the 5,019 normed words, their statistics, and thousands of independently normed rhyme, stem, and fragment cues. Other files provide n x n associative networks for more than 4,000 words and a list of idiosyncratic responses for each normed word. The database will be useful for investigators interested in cuing, priming, recognition, network theory, linguistics, and implicit testing applications. They also will be useful for evaluating the predictive value of free association probabilities as compared with other measures, such as similarity ratings and co-occurrence norms. Of several procedures for measuring preexisting strength between two words, the best remains to be determined. The norms may be downloaded from www.psychonomic.org/archive/.
The aim of the present study was to provide French normative data for 112 action line drawings. The set of action pictures consisted of 71 drawings taken from Masterson and Druks (1998) and 41 additional drawings. It was standardized on six psycholinguistic variables--that is, name agreement, image agreement, image variability, visual complexity, conceptual familiarity, and age of acquisition (AoA). Naming latencies to the action pictures were collected, and a regression analysis was performed on the naming latencies, with the standardized variables, as well as with word frequency and length, taken as predictors. A reliable influence of AoA, name agreement, and image agreement on the naming latencies was observed. The findings are consistent with previous published studies in other languages. The full set of these norms may be downloaded from www.psychonomic.org/archive/.
We provide imageability estimates for 3,000 disyllabic words (as supplementary materials that may be downloaded with the article from www.springerlink.com ). Imageability is a widely studied lexical variable believed to influence semantic and memory processes (see, e.g., Paivio, 1971). In addition, imageability influences basic word recognition processes (Plaut, McClelland, Seidenberg, {\&} Patterson, 1996). In fact, neuroimaging studies have suggested that reading high- and low-imageable words elicits distinct neural activation patterns for the two types e.g., Bedny {\&} Thompson-Schill (Brain and Language 98:127-139, 2006; Graves, Binder, Desai, Conant, {\&} Seidenberg NeuroImage 53:638-646, 2010). Despite the usefulness of this variable, imageability estimates have not been available for large sets of words. Furthermore, recent megastudies of word processing e.g., Balota et al. (Behavior Research Methods 39:445-459, 2007) have expanded the number of words that interested researchers can select according to other lexical characteristics (e.g., average naming latencies, lexical decision times, etc.). However, the dearth of imageability estimates (as well as those of other lexical characteristics) limits the items that researchers can include in their experiments. Thus, these imageability estimates for disyllabic words expand the number of words available for investigations of word processing, which should be useful for researchers interested in the influences of imageability both as an input and as an outcome variable.
The present study introduces the first substantial German database with norms for semantic typicality, age of acquisition, and concept familiarity for 824 exemplars of 11 semantic categories, including four natural (ANIMALS, BIRDS, FRUITS,: and VEGETABLES: ) and five man-made (CLOTHING, FURNITURE, VEHICLES, TOOLS: , and MUSICAL INSTRUMENTS: ) categories, as well as PROFESSIONS: and SPORTS: . Each category exemplar in the database was collected empirically in an exemplar generation study. For each category exemplar, norms for semantic typicality, estimated age of acquisition, and concept familiarity were gathered in three different rating studies. Reliability data and additional analyses on effects of semantic category and intercorrelations between age of acquisition, semantic typicality, concept familiarity, word length, and word frequency are provided. Overall, the data show high inter- and intrastudy reliabilities, providing a new resource tool for designing experiments with German word materials. The full database is available in the supplementary material of this file and also at www.psychonomic.org/archive .
The HAL (hyperspace analog to language) model of lexical semantics uses global word co-occurrence from a large corpus of text to calculate the distance between words in co-occurrence space. We have implemented a system called HiDEx (High Dimensional Explorer) that extends HAL in two ways: It removes unwanted influence of orthographic frequency from the measures of distance, and it finds the $\backslash$nnumber of words within a certain distance of the word of interest (NCount, the number of neighbors). These two changes to the HAL model produce measures of word neighborhood density that are reliably predictive of human lexical decision reaction times.
It is well known that the statistical characteristics of a language, such as word frequency or the consistency of the relationships between orthography and phonology, influence literacy acquisition. Accordingly, linguistic databases play a central role by compiling quantitative and objective estimates about the principal variables that affect reading and writing acquisition. We describe a new set of Web-accessible databases of French orthography whose main characteristic is that they are based on frequency analyses of words occurring in reading books used in the elementary school grades. Quantitative estimates were made for several infralexical variables (syllable, grapheme-to-phoneme mappings, bigrams) and lexical variables (lexical neighborhood, homophony and homography). These analyses should permit quantitative descriptions of the written language in beginning readers, the manipulation and control of variables based on objective data in empirical studies, and the development of instructional methods in keeping with the distributional characteristics of the orthography.
We describe a Windows program that enables users to obtain a broad range of statistics concerning the properties of word and nonword stimuli in an agglutinative language (Basque), including measures of word frequency (at the whole-word and lemma levels), bigram and biphone frequency, orthographic similarity, orthographic and phonological structure, and syllable-based measures. It is designed for use by researchers in psycholinguistics, particularly those concerned with recognition of isolated words and morphology. In addition to providing standard orthographic and phonological neighborhood measures, the program can be used to obtain information about other forms of orthographic similarity, such as transposed-letter similarity and embedded-word similarity. It is available free of charge from www .uv.es/mperea/E-Hitz.zip.
In this study, we compared four expert graders with latent semantic analysis (LSA) to assess short summaries of an expository text. As is well known, there are technical difficulties for LSA to establish a good semantic representation when analyzing short texts. In order to improve the reliability of LSA relative to human graders, we analyzed three new algorithms by two holistic methods used in previous research (Le{\'{o}}n, Olmos, Escudero, Ca{\~{n}}as, {\&} Salmer{\'{o}}n, 2006). The three new algorithms were (1) the semantic common network algorithm, an adaptation of an algorithm proposed by W. Kintsch (2001, 2002) with respect to LSA as a dynamic model of semantic representation; (2) a best-dimension reduction measure of the latent semantic space, selecting those dimensions that best contribute to improving the LSA assessment of summaries (Hu, Cai, Wiemer-Hastings, Graesser, {\&} McNamara, 2007); and (3) the Euclidean distance measure, used by Rehder et al. (1998), which incorporates at the same time vector length and the cosine measures. A total of 192 Spanish middle-grade students and 6 experts took part in this study. They read an expository text and produced a short summary. Results showed significantly higher reliability of LSA as a computerized assessment tool for expository text when it used a best-dimension algorithm rather than a standard LSA algorithm. The semantic common network algorithm also showed promising results.
It has been demonstrated previously that, for some experimental paradigms, Web-based research can reliably replicate lab-based results. Yet questions remain as to what types of research can be reproduced, and where differences arise when they cannot be. The present article examines the effect of research location (laboratory vs. online) on normative data collection tasks. Specifically, participants were randomly assigned to a laboratory or online condition and were asked to rate 593 photorealistic images on the basis of object familiarity (N=103) and object visual complexity (N=98). Dependent measures were compared across location conditions, including response latencies and image rating agreement. Our results suggest that norming data collected online are reliable, but an interesting interplay between task type and research location was observed. Specifically, we found that participating online (i.e., a more familiar environment) leads to systematically higher familiarity ratings than in the lab (i.e., an unfamiliar environment). These differences are not found when the alternate complexity rating task is used.
The use of computer tools has led to major advances in the study of spoken language corpora. One area that has shown particular progress is the study of child language development. Although it is now easy to lexically tag every word in a spoken language corpus, one still has to choose between numerous ambiguous forms, especially with languages such as French or English, where more than 70{\%} of words are ambiguous. Computational linguistics can now provide a fully automatic disambiguation of lexical tags. The tool presented here (POST) can tag and disambiguate a large text in a few seconds. This tool complements systems dealing with language transcription and suggests further theoretical developments in the assessment of the status of morphosyntax in spoken language corpora. The program currently works for French and English, but it can be easily adapted for use with other languages. The analysis and computation of a corpus produced by normal French children 2-4 years of age, as well as of a sample corpus produced by French SLI children, are given as examples.
In this article, we present a set of 12 norms that characterize emotional terms in French, English, German, Spanish, Italian, and Finnish. The high correlation between the norm values in the two emotional dimensions of valence and arousal suggests an interlingual homogeneity of emotional representations and allows a significant metanorm-EMONORM-to be established with 6,383 terms characterized in valence and 4,345 terms characterized in arousal. This metanorm is a resource for creating experimental materials in studies on language and emotions. Furthermore, we perform three tests using EMONORM, with the objectives of (1) identifying basic emotions from their valence and arousal values, (2) determining the orientation of texts referring to positive and negative emotions, and (3) evaluating the intensity of emotions expressed in texts. The results are highly similar to those for human judgments. Finally, we present EMOVAL/SEMOTEX, a Web application for static and dynamic valence and arousal emotional analysis of texts using EMONORM ( http://www.semotex.fr ).