1358 norm sets
The aim of the present study was to provide French normative data for 112 action line drawings. The set of action pictures consisted of 71 drawings taken from Masterson and Druks (1998) and 41 additional drawings. It was standardized on six psycholinguistic variables--that is, name agreement, image agreement, image variability, visual complexity, conceptual familiarity, and age of acquisition (AoA). Naming latencies to the action pictures were collected, and a regression analysis was performed on the naming latencies, with the standardized variables, as well as with word frequency and length, taken as predictors. A reliable influence of AoA, name agreement, and image agreement on the naming latencies was observed. The findings are consistent with previous published studies in other languages. The full set of these norms may be downloaded from www.psychonomic.org/archive/.
Preexisting word knowledge is accessed in many cognitive tasks, and this article offers a means for indexing this knowledge so that it can be manipulated or controlled. We offer free association data for 72,000 word pairs, along with over a million entries of related data, such as forward and backward strength, number of competing associates, and printed frequency. A separate file contains the 5,019 normed words, their statistics, and thousands of independently normed rhyme, stem, and fragment cues. Other files provide n x n associative networks for more than 4,000 words and a list of idiosyncratic responses for each normed word. The database will be useful for investigators interested in cuing, priming, recognition, network theory, linguistics, and implicit testing applications. They also will be useful for evaluating the predictive value of free association probabilities as compared with other measures, such as similarity ratings and co-occurrence norms. Of several procedures for measuring preexisting strength between two words, the best remains to be determined. The norms may be downloaded from www.psychonomic.org/archive/.
We provide imageability estimates for 3,000 disyllabic words (as supplementary materials that may be downloaded with the article from www.springerlink.com ). Imageability is a widely studied lexical variable believed to influence semantic and memory processes (see, e.g., Paivio, 1971). In addition, imageability influences basic word recognition processes (Plaut, McClelland, Seidenberg, {\&} Patterson, 1996). In fact, neuroimaging studies have suggested that reading high- and low-imageable words elicits distinct neural activation patterns for the two types e.g., Bedny {\&} Thompson-Schill (Brain and Language 98:127-139, 2006; Graves, Binder, Desai, Conant, {\&} Seidenberg NeuroImage 53:638-646, 2010). Despite the usefulness of this variable, imageability estimates have not been available for large sets of words. Furthermore, recent megastudies of word processing e.g., Balota et al. (Behavior Research Methods 39:445-459, 2007) have expanded the number of words that interested researchers can select according to other lexical characteristics (e.g., average naming latencies, lexical decision times, etc.). However, the dearth of imageability estimates (as well as those of other lexical characteristics) limits the items that researchers can include in their experiments. Thus, these imageability estimates for disyllabic words expand the number of words available for investigations of word processing, which should be useful for researchers interested in the influences of imageability both as an input and as an outcome variable.
The present study introduces the first substantial German database with norms for semantic typicality, age of acquisition, and concept familiarity for 824 exemplars of 11 semantic categories, including four natural (ANIMALS, BIRDS, FRUITS,: and VEGETABLES: ) and five man-made (CLOTHING, FURNITURE, VEHICLES, TOOLS: , and MUSICAL INSTRUMENTS: ) categories, as well as PROFESSIONS: and SPORTS: . Each category exemplar in the database was collected empirically in an exemplar generation study. For each category exemplar, norms for semantic typicality, estimated age of acquisition, and concept familiarity were gathered in three different rating studies. Reliability data and additional analyses on effects of semantic category and intercorrelations between age of acquisition, semantic typicality, concept familiarity, word length, and word frequency are provided. Overall, the data show high inter- and intrastudy reliabilities, providing a new resource tool for designing experiments with German word materials. The full database is available in the supplementary material of this file and also at www.psychonomic.org/archive .
The HAL (hyperspace analog to language) model of lexical semantics uses global word co-occurrence from a large corpus of text to calculate the distance between words in co-occurrence space. We have implemented a system called HiDEx (High Dimensional Explorer) that extends HAL in two ways: It removes unwanted influence of orthographic frequency from the measures of distance, and it finds the $\backslash$nnumber of words within a certain distance of the word of interest (NCount, the number of neighbors). These two changes to the HAL model produce measures of word neighborhood density that are reliably predictive of human lexical decision reaction times.
It is well known that the statistical characteristics of a language, such as word frequency or the consistency of the relationships between orthography and phonology, influence literacy acquisition. Accordingly, linguistic databases play a central role by compiling quantitative and objective estimates about the principal variables that affect reading and writing acquisition. We describe a new set of Web-accessible databases of French orthography whose main characteristic is that they are based on frequency analyses of words occurring in reading books used in the elementary school grades. Quantitative estimates were made for several infralexical variables (syllable, grapheme-to-phoneme mappings, bigrams) and lexical variables (lexical neighborhood, homophony and homography). These analyses should permit quantitative descriptions of the written language in beginning readers, the manipulation and control of variables based on objective data in empirical studies, and the development of instructional methods in keeping with the distributional characteristics of the orthography.
We describe a Windows program that enables users to obtain a broad range of statistics concerning the properties of word and nonword stimuli in an agglutinative language (Basque), including measures of word frequency (at the whole-word and lemma levels), bigram and biphone frequency, orthographic similarity, orthographic and phonological structure, and syllable-based measures. It is designed for use by researchers in psycholinguistics, particularly those concerned with recognition of isolated words and morphology. In addition to providing standard orthographic and phonological neighborhood measures, the program can be used to obtain information about other forms of orthographic similarity, such as transposed-letter similarity and embedded-word similarity. It is available free of charge from www .uv.es/mperea/E-Hitz.zip.
In this study, we compared four expert graders with latent semantic analysis (LSA) to assess short summaries of an expository text. As is well known, there are technical difficulties for LSA to establish a good semantic representation when analyzing short texts. In order to improve the reliability of LSA relative to human graders, we analyzed three new algorithms by two holistic methods used in previous research (Le{\'{o}}n, Olmos, Escudero, Ca{\~{n}}as, {\&} Salmer{\'{o}}n, 2006). The three new algorithms were (1) the semantic common network algorithm, an adaptation of an algorithm proposed by W. Kintsch (2001, 2002) with respect to LSA as a dynamic model of semantic representation; (2) a best-dimension reduction measure of the latent semantic space, selecting those dimensions that best contribute to improving the LSA assessment of summaries (Hu, Cai, Wiemer-Hastings, Graesser, {\&} McNamara, 2007); and (3) the Euclidean distance measure, used by Rehder et al. (1998), which incorporates at the same time vector length and the cosine measures. A total of 192 Spanish middle-grade students and 6 experts took part in this study. They read an expository text and produced a short summary. Results showed significantly higher reliability of LSA as a computerized assessment tool for expository text when it used a best-dimension algorithm rather than a standard LSA algorithm. The semantic common network algorithm also showed promising results.
It has been demonstrated previously that, for some experimental paradigms, Web-based research can reliably replicate lab-based results. Yet questions remain as to what types of research can be reproduced, and where differences arise when they cannot be. The present article examines the effect of research location (laboratory vs. online) on normative data collection tasks. Specifically, participants were randomly assigned to a laboratory or online condition and were asked to rate 593 photorealistic images on the basis of object familiarity (N=103) and object visual complexity (N=98). Dependent measures were compared across location conditions, including response latencies and image rating agreement. Our results suggest that norming data collected online are reliable, but an interesting interplay between task type and research location was observed. Specifically, we found that participating online (i.e., a more familiar environment) leads to systematically higher familiarity ratings than in the lab (i.e., an unfamiliar environment). These differences are not found when the alternate complexity rating task is used.
The use of computer tools has led to major advances in the study of spoken language corpora. One area that has shown particular progress is the study of child language development. Although it is now easy to lexically tag every word in a spoken language corpus, one still has to choose between numerous ambiguous forms, especially with languages such as French or English, where more than 70{\%} of words are ambiguous. Computational linguistics can now provide a fully automatic disambiguation of lexical tags. The tool presented here (POST) can tag and disambiguate a large text in a few seconds. This tool complements systems dealing with language transcription and suggests further theoretical developments in the assessment of the status of morphosyntax in spoken language corpora. The program currently works for French and English, but it can be easily adapted for use with other languages. The analysis and computation of a corpus produced by normal French children 2-4 years of age, as well as of a sample corpus produced by French SLI children, are given as examples.
In this article, we present a set of 12 norms that characterize emotional terms in French, English, German, Spanish, Italian, and Finnish. The high correlation between the norm values in the two emotional dimensions of valence and arousal suggests an interlingual homogeneity of emotional representations and allows a significant metanorm-EMONORM-to be established with 6,383 terms characterized in valence and 4,345 terms characterized in arousal. This metanorm is a resource for creating experimental materials in studies on language and emotions. Furthermore, we perform three tests using EMONORM, with the objectives of (1) identifying basic emotions from their valence and arousal values, (2) determining the orientation of texts referring to positive and negative emotions, and (3) evaluating the intensity of emotions expressed in texts. The results are highly similar to those for human judgments. Finally, we present EMOVAL/SEMOTEX, a Web application for static and dynamic valence and arousal emotional analysis of texts using EMONORM ( http://www.semotex.fr ).
This article reports, for the first time, type and token frequencies of tones, onsets, codas, rimes, and syllables of Hong Kong Cantonese. The information is derived from a computerized spoken corpus, the Hong Kong Cantonese adult language corpus (HKCAC; Leung {\&} Law, 2001), consisting of more than 140,000 character-syllable units. Since the HKCAC is based on recordings of connected speech, comparisons are made with respect to the inventories of various phonological units between the HKCAC and standard descriptions of the Cantonese phonological system--in particular, Fok (1974) and Bauer and Benedict (1997). It is hoped that the frequency information presented here will become a valuable tool for future psycholinguistic and linguistic research in this language. The full set of these frequency counts may be downloaded from the Psychonomic Society Web archive at www.psychonomic.org/ archive/.
During the last 20 years, psycholinguistic research has identified many variables that influence reading and spelling processes. We describe a new computerized lexical database, LEXOP, which provides quantitative descriptors about the relations between orthography and phonology for French monosyllabic words. Three main classes of variables are considered: consistency of print-to-sound and sound-to-print associations, frequency of orthography-phonology correspondences, and word neighborhood characteristics.
Picture naming was investigated primarily to determine its dependence on certain imagery-related variables, with a secondary aim of developing a new set of Japanese norms for 360 pictures. Pictures refined from the original Nishimoto, Miyawaki, Ueda, Une, and Takahashi (Behavior Research Methods 37:398-416, 2005) set were used. Naming behaviors were measured using four imagery-related measures (imageability, vividness, image agreement, and image variability) and four conventional measures (naming time, name agreement, familiarity, and age of acquisition), as well as a number of other measures (17 total). A simultaneous multiple regression analysis performed on naming times showed that the most reliable predictor was H, a measure of name diversity; two image-related measures (image agreement and vividness) and age of acquisition also contributed substantially to the prediction of naming times. The accuracy of picture naming (measured as name agreement) was predicted by vividness, age of acquisition, familiarity, and image agreement. This suggests that certain processes involving mental imagery play a role in picture naming. The full set of norms and pictures may be downloaded from http://www.psychonomic.org/archive/ or along with the article from http://www.springerlink.com .
Our purpose in the present study is to provide a normative set of nonsensical pictures known as droodles and to demonstrate the role of semantic comprehension in facilitating recall of pictorial stimuli. The set consists of 98 pairs of droodles. Experiment 1 standardized these pictorial stimuli with respect to several variables, such as appropriateness of verbal labels, relationship between two droodles, and correct recall. Appropriateness of verbal labels was rated higher for pictures presented in pairs than for pictures presented singly. Experiment 2 used the standardized set of droodles in a recall experiment similar to those of Bower, Karlin, and Dueck (1975) and others. As we expected, semantic interpretation can strongly facilitate recall. Multiple regression analysis showed that several measures had significant power of explanation for recall performance. The full set of norms and pictures from this article may be downloaded from http://brm.psychonomic-journals.org/content/supplemental.
This study presents a set of sentence contexts and their cloze probabilities for European Portuguese children and adolescents. Seventy-three sentence contexts (35 low- and 38 high-constraint sentence stems) were presented to 90 children and 102 adolescents. Participants were asked to complete the sentence contexts with the first word that came to mind. For each sentence context, responses were listed and cloze probabilities of the words that were chosen to complete the sentence context were computed. Additionally, idiosyncratic and invalid responses (structural and semantic errors) were analyzed. A high degree of consistency in responses among the two age samples (children and adolescents) was found, along with a decrease of idiosyncratic and invalid responses in older participants. These results shed light on age-related changes in the effects of linguistic context on word production, and also in knowledge's representation. The full set of norms may be downloaded from http://brm.psychonomic-journals.org/content/supplemental.
This article presents MANULEX, a Web-accessible database that provides grade-level word frequency lists of nonlemmatized and lemmatized words (48,886 and 23,812 entries, respectively) computed from the 1.9 million words taken from 54 French elementary school readers. Word frequencies are provided for four levels: first grade (G1), second grade (G2), third to fifth grades (G3-5), and all grades (G1-5). The frequencies were computed following the methods described by Carroll, Davies, and Richman (1971) and Zeno, Ivenz, Millard, and Duvvuri (1995), with four statistics at each level (F, overall word frequency; D, index of dispersion across the selected readers; U, estimated frequency per million words; and SFI, standard frequency index). The database also provides the number of letters in the word and syntactic category information. MANULEX is intended to be a useful tool for studying language development through the selection of stimuli based on precise frequency norms. Researchers in artificial intelligence can also use it as a source of information on natural language processing to simulate written language acquisition in children. Finally, it may serve an educational purpose by providing basic vocabulary lists.
This article presents a computerized database of words for use in experimental research in cognitive psychology and psycholinguistics. The data are based on the oral vocabulary of 200 Spanish-speaking children aged from 11.16 to 49.16 months. The database includes 15,428 Spanish words (tokens) and comprises 1,259 different words (types). It provides information about age of acquisition, orthography, grammar, semantics, and frequency.
Several methods to study the recognition and similarity of alphanumeric characters are briefly discussed and evaluated. In particular, the application of the choice-model (Luce, 1959, 1963) to recognition of letters is criticized. A feature analytic model for recognition of alphanumeric characters based on Tversky's (1977) features of similarity is proposed and tested. It is argued that the proposed model: (a) is parsimonious in that it utilizes a relatively small number of parameters, (b) is psychologically more meaningful compared with other approaches in that it is attempting to study underlying processes rather than just reveal a similarity structure, (c) yields predictions that have a high level of fit with the observed data. Possible implications from the use of the model for future research are briefly discussed.
The confusion matrix for the full lowercase English alphabet is estimated, based upon 25 trials per letter for each of seven subjects. Average correct recognition was controlled to 0.5 by limiting brightness and duration of displays to individually determined levels. Comparison of the obtained data to that reported by Bouma 11971t for eccentric vision supports the conclusion that limited energy foveal recognition is qualitatively different from eccentric vision recognition. Comparison of the obtained data to the uppercase confusion matrix reported by Townsend (1971) supports the inference that recognition performance has more between-letter variability for both recognizability and confusion pairings for the lowercase alphabet than for the upper.
A number of inconsistencies are evident in the literature examining word-neighborhood size and frequency effects. One reason for the inconsistency may be that there are no standardized materials and criteria used in the different studies. Each experimenter has devised his or her word neighborhoods using different criteria for neighborhood size and frequency. The purpose of the present study was to develop a standardized set of word neighborhoods. 800 orthographic neighborhoods were constructed with 4- and 5-letter words. The word lists were devised relative to the key elements that have been identified in the literature: (1) target-word frequency, (2) number of words in the neighborhood, (3) number of words higher in frequency than the target word, (4) number of letter positions contributing to the neighborhood, and (5) summation of the frequency of all neighbors (providing a standard metric for high- vs low-frequency neighborhoods). ((c) 1998 APA/PsycINFO, all rights reserved)
A comprehensive count of bigram frequencies and versatilities by position was tabulated for 2-9 letter words recorded by H. Kucera and W. Francis (1967). A total of 577 bigrams were found variously distributed throughout words. Such counts should prove useful in determining the orthographic regularity of specific words. (10 ref) (PsycINFO Database Record (c) 2006 APA, all rights reserved)
Tabulations of letter and letter-combination versatility and frequency were made based on the Kucera and Francis (1967) word frequency count. Letter versatility, a new descriptive statistic, was defined as the number of different words in which a letter appears. These tabulations may be useful in the investigation of visual information processing, reading skills, and human memory.
Researchers often require subjects to make judgments that call upon their knowledge of the orthographic structure of English words. Such knowledge is relevant in experiments on, for example, reading, lexical decision, and anagram solution. One common measure of orthographic structure is the sum of the frequencies of consecutive bigrams in the word. Traditionally, researchers have relied on token-based norms of bigram frequencies. These norms confound bigram frequency with word frequency because each instance (i.e., token) of a particular word in a corpus of running text increments the frequencies of the bigrams that it contains. In this article, the authors report a set of type-based bigram frequencies in which each word (i.e., type) contributes only once, thereby unconfounding bigram frequency from word frequency. The authors show that type-based bigram frequency is a better predictor of the difficulty of anagram solution than is token-based frequency. These norms can be downloaded from www.psychonomic.org/archive/ .
CVCVC (319) words and paralogs previously assessed for associative reaction time (RT) by Taylor and Kimble were assessed for rated frequency (a′) and scaled rated meaningfulness (m′) following procedures used by Noble. Reliability of the a′ scale, based on three intergroup correlations, resulted in rs of .92, .90 and .89. Reliability of the m′ scale, based on an internal consistency test resulted in a mean discrepancy between the 128 empirical proportions and their corresponding theoretical proportions of 2.4; that is, the average error of reproducing all the original data from m′ scaled values was 2.4{\%}. The r between m′ and RT was .71. {\textcopyright} 1972 Academic Press, Inc. All Rights Reserved.
To understand how and when object knowledge influences the neural underpinnings of language compre- hension and linguistic behavior, it is critical to determine the specific kinds of knowledge that people have. To extend the normative data currently available, we report a relatively more comprehensive set of object attribute rating norms for 559 concrete object nouns, each rated on seven attributes corresponding to sensory and motor modalities—color, mo- tion, sound, smell, taste, graspability, and pain—in addition to familiarity (376 raters, M 0 23 raters per item). The mean ratings were subjected to principal-components analysis, revealing two primary dimensions plausibly interpreted as relating to survival. We demonstrate the utility of these ratings in accounting for lexical and semantic decision latencies. These ratings should prove useful for the design and interpretation of experimental tests of conceptual and perceptual object processing.
We tabulated upper- and lowercase letter frequency using several large-scale English corpora (approximately 183 million words in total). The results indicate that the relative frequencies for upper- and lowercase letters are not equivalent. We report a letter-naming experiment in which uppercase frequency predicted response time to uppercase letters better than did lowercase frequency. Tables of case-sensitive letter and bigram frequency are provided, including common nonalphabetic characters. Because subjects are sensitive to frequency relationships among letters, we recommend that experimenters use case-sensitive counts when constructing stimuli from letters.
Notes that anagram solution time may be affected by bigrams in the solution word. A data base is presented from which solution anagrams may be created based on bigram frequency and versatility.
372 CLARK AND PAIVIO Shoben, 1983), and as theoretical critics questioned the proposed nature of imagery (e.g., Pylyshyn, 1973). Research on item attributes continues to become more sophisticated in a variety of ways, including considera-tion of larger numbers of properties (e.g., Paivio et al., 1989; Rubin, 1980), and serious efforts to model and simulate the effects of stimulus attributes (e.g., Ellis {\&} Lambon Ralph, 2000). Computers have played a central role in both of these developments. The multivariate and simulation studies that might ultimately contribute to the emergence of theoretical models with strong empirical foundations has increased demand for large item pools and increased numbers of properties. The need for a greater number of properties and items follows from the essentially nonexperimental nature of item attribute research. Although researchers may exper-imentally assign different types of materials to different subjects, they generally do not and in some cases cannot experimentally manipulate the property or properties of interest (although such experimental approaches have been used in some cases). Rather, the properties are gen-erally measured in some fashion, and these measures will invariably correlate with other properties that might produce spurious effects or mask the effects of the target attributes. The primary way to address this problem, as in any nonexperimental research, is to identify diverse potentially contaminating constructs, obtain reliable and valid measures, and either control them in the selection of materials or include them in statistical analyses that can accommodate correlated factors (e.g., multiple re-gression, factor analysis, structural equation modeling). In addition to serving this control function, the col-lection of a large number of properties can itself provide information useful in the conceptualization of various item attributes. We can illustrate with one controversial question—the relative importance of frequency and age of acquisition in picture naming and other semantic re-trieval tasks (Morrison {\&} Ellis, 2000). Paivio et al. (1989) obtained picture naming and imagery latencies for a moderate-sized pool of pictures and their most common labels. Information on a wide range of properties, in-cluding age of acquisition, was also obtained. Factor analysis of the results indicated that age of acquisition loaded on several different factors (e.g., familiarity, con-creteness, name length), all of which contributed to pic-ture naming latencies. One interpretation of this result is that age of acquisition is a multidimensional measure that taps a number of distinct properties of words and pictures, hence its superiority to single-component pre-dictors in multiple regression analyses. More specula-tively, one might hypothesize that people rating age of acquisition are actually making judgments of how con-crete, short, and familiar items are, and that children in fact first learn words that tend to be concrete, short, and familiar. Generalization across items provides yet another rea-son to continue the development of item norms for use in cognitive research and theorizing. No single set of norms will ever suffice, because results can depend on the par-ticular pool of items that have been included in the norms. Despite the large corpus of materials on which the Ku{\v{c}}era and Francis (1967) frequency norms are based, for exam-ple, abstract words are still probably overrepresented just because of the types of text that dominate the corpus (e.g., literary and academic materials). It is therefore im-portant to continue to develop additional norms to per-mit evaluation of the generality of findings across di-verse word pools. Another facet of the generalization issue is the possi-bility of generational or cohort differences across ex-tended periods of time. With respect to word familiarity, for example, exposure to and knowledge of particular words might differ today from ratings, like those in the PYM norms, collected during the 1960s. New norms and replication of existing properties allow researchers to de-termine the continuing validity of norms collected years and in some cases decades ago. This article reports two extensions of the PYM norms. Part 1 reports a marked expansion of the number of prop-erties available for the original 925 PYM items, and Part 2 reports an expansion of the number of items for which basic properties are available. We also provide re-sults of factor analyses for both extensions, with the analysis in Part 1 being particularly informative about interrelationships among a diverse collection of word properties.
Ratings of familiarity and pronounceability were obtained for a sample of 199 names and 199 nouns. Frequency and familiarity were more closely related in the proper name pool than the word pool, although the correlation was modest in both cases. Familiarity and pronounceability were highly related for both names and nouns.
We collected imageability and body-object interaction (BOI) ratings for 599 multisyllabic nouns. We then examined the effects of these variables on a subset of these items in picture-naming, word-naming, lexical decision, and semantic categorization. Picture-naming latencies were taken from the International Picture-Naming Project database (Szekely, Jacobsen, D'Amico, Devescovi, Andonova, Herron, et al. Journal of Memory and Language, 51, 247-250, 2004), word-naming and lexical decision latencies were taken from the English Lexicon Project database (Balota, Yap, Cortese, Hutchison, Kessler, Loftis, et al. Behavior Research Methods, 39, 445-459, 2007), and we collected semantic categorization latencies. Results from hierarchical multiple regression analyses showed that imageability and BOI separately accounted for unique latency variability in each task, even with several other predictor variables (e.g., print frequency, number of syllables and morphemes, age of acquisition) entered first in the analyses. These ratings should be useful to researchers interested in manipulating or controlling for the effects of imageability and BOI for multisyllabic stimuli in lexical and semantic tasks.
Used 2 techniques to gather data on property dominance and property goodness. Sensory properties of verbally depicted items, such as those of color or shape, were indexed for dominance (frequency of output) and typicality (perceptual goodness). 193 college students were each assigned to 1 of 4 conditions. The most dominant property response was computed for 105 nouns by 1 group. A 2nd group rated these properties for typicality, relative to a constituent property (i.e., given one's idea of yellow, how typically "yellow" were specific items?). A 3rd group produced as many properties as possible for each of 65 nouns. Dominance was computed for all 459 properties so produced. The 4th group rated these 459 properties for typicality relative to the parent noun. A multimethod-multitrait analysis indicated that both typicality and dominance were reliable and that both exhibited convergent and discriminant validity. Typicality measured relative to an ideal property exhibited greater discriminant validity from dominance than when measured relative to the parent noun. Selected uses of these norms in studies of semantic memory, metaphor judgment, and concept identification are discussed. (PsycINFO Database Copyright 1983 American Psychological Assn, all rights reserved)
To assist research in anagram solving, this paper presents bigram statistics for 205 5-letter words known to form single-solution anagrams. Imagery, concreteness, age-of-acquisition, familiarity, and meaningfulness values for these words have previously been published. The bigram statistics presented here include the bigram rank and GTZERO (a variable involving the number of nonzero entries in the bigram frequency matrix) measures recently devised and tested by G. A. Mendelsohn (1976). ((c) 1999 APA/PsycINFO, all rights reserved)
Imageability ratings made on a 1-7 scale and reaction times for 3,000 monosyllabic words were obtained from 31 participants. Analyses comparing these ratings to 1,153 common words from Toglia and Battig (1978) indicate that these ratings are valid. Reliability was assessed (alpha = .95). The information obtained in this study adds to that of other normative studies and is useful to researchers interested in manipulating or controlling imageability in word recognition and memory studies. These norms can be downloaded from www.psychonomic.org/archive/.
Although the typicality effect has been much studied in the semantic memory literature, typ-icality ratings exist for exemplars from only a very limited number of categories. This lack of ratings frequently limits the range of stimuli that can be used in investigations of the typicality effect. As an aid in stimulus construction, this paper reports typicality ratings of 893 exemplars from 93 different categories.
The positional frequency and versatility of letters were tabulated for six-, seven-, and eight-letter English words.
Examined whether a rating-based procedure that has already been used by other investigators can be used for derivation of typicality ratings from children. In Exp 1, 96 kindergartners (aged 4.3-6 yrs) generated typicality for items belonging to each of 4 categories. The same Ss participated in Exp 2, in which they were asked to generate attributes for the members of 4 categories for the benefit of 2 people from another planet who had no knowledge about the particular items. In Exp 3, with 120 university students (aged 19-21.3 yrs), the correspondence between adults' family resemblance and typicality ratings on the same materials was tested. Results show that the procedure cannot be reliably used for this purpose; children rated category items in terms of personal preferences rather than as a function of how representative they considered the items to be of their superordinate category. On the basis of these findings, the authors propose an alternative method based on the family resemblance scores of the category members to derive typicality ratings from young children. (PsycINFO Database Record (c) 2012 APA, all rights reserved)
The extent to which an item is a prototypical exemplar of a category has been found to predict several experimental results (e.g.,reaction times in category classification, free and cued recall of lists, release from proactive inhibition in recall). We present prototypicality ratings for 840 words, equally distributed over 28 categories. The categories were taken from Battig and Montague's (1969) normative tables; only those categories that contained "concrete" items in common usage were employed in the study. Intragroup reliability correlations were high for all categories tested, as were the correlations for prototypicality ratings between the present study and that of Rosch (1975). In addition, correlations between prototypicality ratings, production frequencies, and word frequencies of the items are given.
Six groups of subjects rated word pairs for the degree to which they exemplified one of six semantic relationships. The relationships that subjects were instructed to rate were antonymy, synonymity, subordination, superordination, coordination, and similarity. Stimulus pairs represented antonyms, synonyms, subordinates, superordinates, and coordinates. The pairs representing each stimulus relationship varied across four levels of typicality, ranging from good examples of the relationship to unrelated pairs. The highest rating in each group was given to the stimulus relationship corresponding to the relationship being judged (e.g., antonyms received the highest rating under antonym judgment instructions). This interaction was strongest for high-typicality pairs and decreased across the levels of typicality. Semantic decision models cannot explain these results unless the models are modified so that decisions are based on relationship similarity, the degree to which a stimulus pair exemplifies the relationship subjects are instructed to judge.
Three hundred and seventy-five children in Grades 2, 3, 4, and 6 were asked to generate instances of 25 different categories within a time period of 1 min per category. Data were tallied so that category instances are ranked as to proportion of subjects making each response at each grade level. Indications of the average number of category instances generated by children at each grade level within each category are provided.
Homophones are words that share phonology but differ in meaning and spelling (e.g., beach, beech). This article presents the results of normative surveys that asked young and older adults to free associate to and rate the dominance of 197 homophones. Although norms exist for young adults on word familiarity and frequency for homophones, these results supplement the literature by (1) reporting the four most frequent responses to visually presented homophones for both young and older adults, and (2) reporting young and older adults' ratings of homophone dominance. Results indicated that young and older adults gave the same first response to 67{\%} of the homophones and rated homophone dominance similarly on 60{\%} of the homophone sets. These results identify a subset of homophones that are preferable for research with young and older adults because of age-related equivalence in free association and dominance ratings. These norms can be downloaded from the Psychonomic Society's Web archive, www.psychonomic.org/archive/.
First- and second-order approximations to English and orthographic neighbor ratio values are provided for Paivio, Yuille, and Madigan's (1968) 925 nouns. First- and second-order approximations to English are information theory measures of the probability of generating a word on a letter-by-letter basis. The orthographic neighbor ratio is the frequency of a word divided by the sum of the frequencies of all words that can be generated by changing one of its letters. Thus, the orthographic neighbor ratio provides a measure of a sophisticated guessing model in which partial information about a word is obtained and a decision is made on the basis of the relative frequencies of the possible responses. Correlations with existing norms are reported.
Mean probabilities of acoustic confusability were computed for 1172 CCC trigrams chosen at random from Witmer?s list of association values, and the results were tabulated. ? 1969, Psychonomic Journals, Inc.. All rights reserved.
Words that are homonyms-that is, for which a single written and spoken form is associated with multiple, unrelated interpretations, such as COMPOUND, which can denote an {\textless} enclosure {\textgreater} or a {\textless} composite {\textgreater} meaning-are an invaluable class of items for studying word and discourse comprehension. When using homonyms as stimuli, it is critical to control for the relative frequencies of each interpretation, because this variable can drastically alter the empirical effects of homonymy. Currently, the standard method for estimating these frequencies is based on the classification of free associates generated for a homonym, but this approach is both assumption-laden and resource-demanding. Here, we outline an alternative norming methodology based on explicit ratings of the relative meaning frequencies of dictionary definitions. To evaluate this method, we collected and analyzed data in a norming study involving 544 English homonyms, using the eDom norming software that we developed for this purpose. Dictionary definitions were generally sufficient to exhaustively cover word meanings, and the methods converged on stable norms with fewer data and less effort on the part of the experimenter. The predictive validity of the norms was demonstrated in analyses of lexical decision data from the English Lexicon Project (Balota et al., Behavior Research Methods, 39, 445-459, 2007), and from Armstrong and Plaut (Proceedings of the 33rd Annual Meeting of the Cognitive Science Society, 2223-2228, 2011). On the basis of these results, our norming method obviates relying on the unsubstantiated assumptions involved in estimating relative meaning frequencies on the basis of classification of free associates. Additional details of the norming procedure, the meaning frequency norms, and the source code, standalone binaries, and user manual for the software are available at http://edom.cnbc.cmu.edu .
A list of English palindromes that approaches comprehensiveness is given. A stochastic model for the distribution of palindromes within the language is proposed. The model relates frequencies of palindromes and of heteropalindromes (Jones, 1980) as a function of word length. and it is shown to predict accurately the actual numbers of palindromes that occur
Ratings of orthographic distinctiveness were obtained for 139 homonym pairs. Mean ratings on a 9-point scale ranged from 7.75 to 2.44. Reliability of the ratings was high (r = .91). In addition, orthographic distinctiveness was found to be independent of disparity in perceived meanings of the separate homonym forms.
Over the past two decades, homographs have been used in psychological experiments aimed at testing a variety of theoretical issues concerning memory and language. Often, such research requires prior knowledge of the dominance relations among various meanings of the homographs. Previously available homograph meaning norms are limited because they are now more than 10 years old, and they have typically reported only the two most dominant meanings even though many homographs have three or more common meanings. This paper presents normative data on 120 homographs from a relatively large, heterogeneous sample of subjects (N = 100). Meaning dominance was assessed by having subjects write the first definition that came to mind for each homograph. Definition responses were grouped by similarity, and the resulting meaning categories were verified against dictionary meaning classifications. The number of distinct meanings varied from two to six for the homographs investigated, and frequency of response is reported for all definition categories. {\textcopyright} 1994 Psychonomic Society, Inc.
Category typicality norms from 12 natural language categories are presented for kindergarten, third-grade, sixth-grade, and college students. Subjects first selected examples of familiar word concepts and rated them on a 3-point scale in terms of category typicality. Age differences in the percentage of items included as category members were found primarily for the less typical items, with inclusion rates varying as a function of both age and typicality level. The absolute level of typicality judgments increased with age, although correlations between the children's and college students' ratings were generally significant for all three children's groups, with average correlations increasing somewhat with age. It was suggested that the rating data would be useful to developmental investigators interested in children's processing of category information.
Norms were collected to determine the relative dominance of different meanings of homo- graphic words. Forty-six subjects wrote down the first word that came to mind for each of 320 homographs. Each homograph, the number of times each meaning was given, and the specific associates are made available. In addition, correlations with other norms are presented.
Abstract A description is presented of normative data for property responses to 121 words— 17 category labels, three typical and three atypical members of each category , and the words “plant” and “animal.” The production frequency of properties is considered a ... $\backslash$n