1396 norm sets
It is generally assumed that orthographic-phonological (O-P) consistencies are higher for Japanese kana words than for kanji words and that orthographic-semantic (O-S) consistencies are higher for kanji words than for kana words. In order to examine the validity of these assumptions, we attempted to measure the O-P and O-S consistencies for 339 kana words and 775 kanji words. Orthographic neighbors were first generated for each of these words. In order to measure the O-P consistencies of the words, their neighbors were then classified as phonological friends or enemies, based on whether the characters shared with the original word were pronounced the same in the two words. In order to measure the O-S consistencies, the similarity in meaning of each of the neighbors to the original word was rated on a 7-point scale. Based on the ratings, the neighbors were classified as semantic friends or enemies. The results indicated that both the O-P consistencies for kanji words and the O-S consistencies for kana words were greater than previously assumed and that the two scripts were actually quite similar on both types of consistency measures. The implications for the nature of the reading processes for kana and kanji words are discussed.
When using verbal stimuli, researchers usually equate words on frequency of use. However, for some ambiguous words (e.g., ball as a round object or a formal dance), frequency counts fail to distinguish how often a particular meaning is used. This study evaluates the use of ratings to estimate meaning frequency. Analyses show that ratings correlate highly with word frequency counts when orthographic and meaning frequencies should converge, are not unduly influenced by semantic factors, and may provide a better measure of relative meaning dominance than the word association task does. Furthermore, the ratings allow researchers to equate or manipulate frequency of meaning use for ambiguous and unambiguous words. Ratings for 211 words are reported.
There exist surprisingly few normative lists of word meanings even though homographs—words having single spellings but two or more distinct meanings—are useful in studying memory and language. The meaning norms that are available all have one or more weaknesses, including: (1) the collection of free associates rather than meanings as responses to the stimulus words; (2) the collection of single rather than multiple responses to the stimulus words; (3) the inclusion of only the two most frequently occurring meaning categories, rather than all meaning categories, for the stimulus words; (4) omission of the responses typical of each meaning category; (5) inadequate randomization of the presentation order of the stimulus words; and (6) unpaced presentation of the stimulus words. We have compiled meaning norms for 90 common English words of low, medium, and high concreteness using a methodology designed to correct these weaknesses. Analysis showed that words of medium concreteness have significantly more first-response meanings than do words of either low or high concreteness, lending support to the view that concreteness is a categorical, rather than a continuous, semantic attribute.
We present a set of 150 pictures with morphologically complex English compound names. The pictures were collected from various sources and were standardized to appear as grayscale line drawings of a fixed size. All the compounds had two constituents and were primarily of the noun-noun type. Following previous studies, we collected name agreement (percentage and H), familiarity, image agreement, and visual complexity norms, as well as frequency estimates for the whole compound word and its first and second constituents. These pictures and their corresponding norms (available from the Psychonomic Society's supplemental archive) are a valuable tool in the study of the morphological representation of complex words in language processing.
Many recent studies have demonstrated the influence of sublexical frequency measures on language processing, or called for controlling sublexical measures when selecting stimulus material for psycholinguistic studies (Aichert {\&} Ziegler, 2005). The present study discusses which measures should be controlled for in what kind of study, and presents orthographic and phonological syllable, dual unit (bigram and biphoneme) and single unit (letter and phoneme) type and token frequency measures derived from the lemma and word form corpora of the CELEX lexical database (Baayen, Piepenbrock, {\&} Gulikers, 1995). Additionally, we present the SUBLEX software as an adaptive tool for calculating sublexical frequency measures and discuss possible future applications. The measures and the software can be downloaded at www.psychonomic.org.
Although word co-occurrences within a document have been demonstrated to be semantically useful, word interactions over a local range have been largely neglected by psychologists due to practical challenges. Shannon's (Bell Systems Technical Journal, 27, 379–423, 623–665, 1948) conceptualization of information theory suggests that these interactions should be useful for understanding communica- tion. Computational advances make an examination of local word–word interactions possible for a large text corpus. We used Brants and Franz's (2006) dataset to generate conditional probabilities for 62,474 word pairs and entropy calculations for 9,917 words in Nelson, McEvoy, and Schreiber's (Behavior Research Methods, Instruments, {\&} Computers, 36, 402–407, 2004) free association norms. Semantic asso- ciativity correlated moderately with the probabilities and was stronger when the two words were not adjacent. The number of semantic associates for a word and the entropy of a word were also correlated. Finally, language entropy decreases from 11 bits for single words to 6 bits per word for four-word sequences. The probabilities and entropies discussed here are included in the supplemental materials for the article.
Early vocabulary development is a reliable predictor of children's later language skills. The MacArthur-Bates Communicative Development Inventory (CDI) has provided a powerful tool to assess earlyvocabulary development in English and other languages. However, there have been no published CDI norms for Mandarin Chinese. Given the importance of large-scale comparative data sets for understanding the early childhood lexicon, we have developed an early vocabulary inventory for Mandarin. In this article, we report our efforts in developing this instrument, and discuss the data collected from 884 Chinese families in Beijing over a period of 12-30 months, based on our instrument. Chinese children's receptive and expressive lexicons as assessed by our inventory match well with those reported for English on the basis of CDI. In particular, our data indicate comprehension-production differences, individual differences in early comprehension and in later production, and different lexical development profiles among infants versus toddlers. We also make the checklists and norms of our inventory available to the research community via the Internet; they may be accessed from the Psychonomic Society's Archive of Norms, Stimuli, and Data, at www.psychonomic.org/archive.
Age of acquisition and imageability ratings were collected for 2,645 words, including 892 verbs and 213 function words. Words that were ambiguous as to grammatical category were disambiguated: Verbs were shown in their infinitival form, and nouns (where appropriate) were preceded by the indefinite article (such as to crack and a crack). Subjects were speakers of British English selected from a wide age range, so that differences in the responses across age groups could be compared. Within the sub- set of early acquired noun/verb homonyms, the verb forms were rated as later acquired than the nouns, and the verb homonyms of high-imageability nouns were rated as significantly less imageable than their noun counterparts. A small number of words received significantly earlier or later age of acquisition rat- ings when the 20–40 years and 50–80 years age groups were compared. These tend to comprise words that have come to be used more frequently in recent years (either through technological advances or so- cial change), or those that have fallen out of common usage. Regression analyses showed that although word length, familiarity, and concreteness make independent contributions to the age of acquisition measure, frequency and imageability are the most important predictors of rated age of acquisition.
Determined month-by-month norms for comprehension and production of 396 words from 8 to 16 mo, and production of 680 words from 16 to 30 mo, derived from a norming study of 1,789 children aged 8-30 mo that used the Communicative Development Inventories. The norms are available in the form of a database program, LEX, for MS-DOS-based computers.
Conducted 7 studies to determine lower bounds for the amount of lexical ambiguity of words used in English text. Two special types of ambiguity are described and quantified, and a refined method for quantifying the ambiguity of individual lexical units is presented. Results of the studies indicate that at least 32{\%} of the words used in English text were ambiguous; it is suggested, however, that this figure is probably conservative. Temporary word definitions established for special purposes occurred in 30{\%} of the sample of texts. (21 ref) (PsycINFO Database Record (c) 2012 APA, all rights reserved)
A corpus of 576 words and orthographically legal pseudowords was rated by 150 undergraduates to obtain a subjective estimate of the number of meanings possessed by the stimuli. The information contained in this corpus may be used to supplement current, sources of word-meaning information (e.g., total number of dictionary entries). Experimental evidence is presented that supports the reliability of the normative data.
Age-of-acquisition, imagery, concreteness, familiarity, and ambiguity measures for 1,944 words of varying length and frequency of occurrence are presented. The words can all be used as nouns. Intergroup reliabilities are satisfactory on all attributes. Correlations with pre-vious word lists are significant, and the intercorrelations between measures match previous findings. METHOD Table 1 Distribution of Words by Length and Frequency in the Sample of 1,944 Drawn From Thorndike-Lorge (1944) many alternative meanings are unknown or of low salience to most subjects. However, words may be expected to vary in their degree of effective ambiguity, and it is this that we set out to measure. L;;.9 Word Length (L) L.;;5 Thorndike-Lorge Frequency Word Sample A total of 1,944 words were selected from the Thorndike-Lorge (1944) word count by means of a semirandom procedure, in such a way as to fill the cells of a 3 by 3 matrix of word length and frequency (Table I). The overall strategy was to select every 10th word that could be used as a noun. If this was not suitable, the first appropriate word within that group of 10 was selected. It was decided that, for research purposes, a fairly even distribution of words over length by frequency combinations was desirable. Due to the underrepresentation of infrequent short words and frequent long words, the latter were selected in preference to other frequency by length combina-tions to give a more even distribution. The numbers of words selected in each cell aregiven in Table I. Ratings Procedure F or the age-of-acquisition, imagery, concreteness, and famil-iarity ratings, the following procedure was used. The words were printed, in random order, 20 to a page, and alongside each word was a 7-point rating scale. The pages were then shuffled and assembled into booklets, so that each booklet contained the pages in a different random order. Due to the large number of pages, the booklets were divided into three sections, each containing approximately 33{\%} of the words. The three
Solution-word imagery appears to affect difficulty of anagram solving. To assist research in this area, imagery, concreteness, age-of-acquisition, familiarity, and meaningfulness values for 205 5-letter words, known to form single-solution anagrams, are presented. None of the words have repeated letters. In an empirical study with university students, intergroup reliabilities were satisfactory on all attributes. Significant correlations were found with previous word lists, and the intercorrelations between dimensions matched previous findings. (PsycINFO Database Record (c) 2008 APA, all rights reserved).
In psychology, lexical norms related to the semantic properties of words, such as concreteness and valence, are important research resources. Collecting such norms by asking judges to rate the words is very time consuming, which strongly limits the number of words that compose them. In the present article, we present a technique for estimating lexical norms based on the latent semantic analysis of a corpus. The analyses conducted emphasize the technique's effectiveness for several semantic dimensions. In addition to the extension of norms, this technique can be used to check human ratings to identify words for which the rating is very different from the corpus-based estimate.
400 undergraduates completed booklets made up of words from the Toronto Word Pool. Ss rated the words for imagery or concreteness. Imagery was defined as the ease with which a word aroused a mental image, and concreteness was defined in relation to level of abstraction. The degree to which a word was functionally a noun was estimated in a sentence generation task. The mean and standard deviation of the imagery and concreteness ratings for each item were derived, together with letter and printed frequency counts for the words and indications of sex differences in the ratings. In a follow-up study with 120 undergraduates, norms included a grammatical function code derived from dictionary definitions, a percent noun judgment, indices of statistical approximation to English, and an orthographic neighbor ratio.
The age at which words are first learned appears to be more influential in determining the ease of retrieving words from semantic memory than objective frequency, familiarity, imagery, and meaningfulness. To facilitate research on a wider variety of tasks, we present norms for 543 words for age-of-acquisition, imagery, familiarity, and meaningfulness. Most of the words form single-solution anagrams. There are 471 six-letter nouns and 72 five-letter words. Also reported are the means, 80s, and ranges for each dimension and the intercorrelations between dimensions. Intergroup reliabilities ranged from .847 to .982. Recent studies have indicated that the age at which words are first learned is influential in determining the ease of retrieving words from semantic memory. Frequency of usage during childhood was found to predict latency to name category instances more accurately than adult frequency of usage (Loftus {\&} Suppes, 1972). Age-of-acquisition as rated by young adults was found by Carroll and White (1973) to be a more relevant variable than objective frequency in predicting latency to name pictures. Rated age -of-acquisition also predicted the speed and likelihood of solving anagrams more accurately than rated familiarity, objective frequency, imagery, and meaningfulness (Stratton, Jacobus, {\&} Brinley, Note l). The norms reported in this paper provide adult norms on rated age-of-acquisition, rated imagery, rated familiarity, and meaningfulness for 543 words. Two word samples from these norms were used in an earlier study (see Note 1). METHOD
In the present study, we presented picture-naming latencies along with ratings for a set of important characteristics of pictures and picture names: age of acquisition, frequency, picture-name agreement, name agreement, visual complexity, familiarity, and word length. The validity of these data was established by calculating correlations with previous studies. Regression analyses show that our ratings account for a larger amount of variance in RTs than do previous data. RTs were predicted by all variables except complexity and length. A complete database presenting details about all of these variables is available in the supplemental materials, downloadable from http://brm.psychonomic-journals.org/content/supplemental.
This paper reports imagery ratings for 338 nouns by a method similar to that of Paivio. For 111 of the nouns a comparison with ratings reported by Paivio with those reported here yielded a correlation of 0.944. Selection of nouns rated for this study was made from stimulus terms employed in various free-word-association studies.
Sensory experience ratings (SERs) reflect the extent to which a word evokes a sensory and/or perceptual experience in the mind of the reader. Juhasz, Yap, Dicke, Taylor, and Gullick (Quarterly Journal of Experimental Psychology 64:1683-1691, 2011) demonstrated that SERs predict a significant amount of variance in lexical-decision response times in two megastudies of lexical processing when a large number of established psycholinguistic variables are controlled for. Here we provide the SERs for the 2,857 monosyllabic words used in the Juhasz et al. study, as well as newly collected ratings on 3,000 disyllabic words. New analyses with the combined set of words confirmed that SERs predict a reliable amount of variance in the lexical-decision response times and naming times from the English Lexicon Project (Balota, Yap, Cortese, Hutchison, Kessler, Loftus, {\&} Treiman, Behavior Research Methods 39:445-459, 2007) when a large number of surface, lexical, and semantic variables are statistically controlled for. The results suggest that the relative availability of sensory/perceptual information associated with a word contributes to lexical-semantic processing.
Normative values on various word characteristics were obtained for abstract, concrete, and emotion words in order to facilitate research on concreteness effects and on the similarities and differences among the three word types. A sample of 78 participants rated abstract, concrete, and emotion words on concreteness, context availability, and imagery scales. Word associations were also gathered for abstract, concrete, and emotion words. The data were used to investigate similarities and differences among these three word types on word attributes, association strengths, and number of associations. These normative data can be used to further research on concreteness effects, word type effects, and word recognition for abstract, concrete, and emotion words.
The problem of word ambiguity has generally been overlooked in compiling lists of words measured on various attributes. In this study, rating measures were obtained on the meanings of 387 words, the ambiguity of which had been established empirically. Imagery, age-of-acquisition, concreteness, and familiarity ratings are reported for each meaning, together with an index of meaning dominance. The results suggest that the most dominant meanings tend to be the most imageable, concrete, familiar, and earliest acquired. Generally satisfactory correlations with other norms were obtained.
We present a new database of Dutch word frequencies based on film and television subtitles, and we validate it with a lexical decision study involving 14,000 monosyllabic and disyllabic Dutch words. The new SUBTLEX frequencies explain up to 10{\%} more variance in accuracies and reaction times (RTs) of the lexical decision task than the existing CELEX word frequency norms, which are based largely on edited texts. As is the case for English, an accessibility measure based on contextual diversity explains more of the variance in accuracy and RT than does the raw frequency of occurrence counts. The database is freely available for research purposes and may be downloaded from the authors' university site at http://crr.ugent.be/subtlex-nl or from http://brm.psychonomic-journals.org/content/supplemental.
To facilitate investigations of verbal emotional processing, we introduce$\backslash$nthe Leipzig Affective Norms for German (LANG), a list of 1,000 German$\backslash$nnouns that have been rated for emotional valence, arousal, and concreteness.$\backslash$nA critical factor regarding the quality of normative word data is$\backslash$ntheir reliability. We therefore acquired ratings from a sample that$\backslash$nwas tested twice, with an interval of 2 years, to calculate test-retest$\backslash$nreliability. Furthermore, we recruited a second sample to test reliability$\backslash$nacross independent samples. The results show (1) the typical quadratic$\backslash$nrelation of valence and arousal, replicating previous data, (2) very$\backslash$nhigh test-retest reliability ({\textgreater}.95), and (3) high correlations between$\backslash$nthe two samples ({\textgreater}.85). Because the range of ratings was also very$\backslash$nhigh, we provide a comprehensive set of words with reliable affective$\backslash$nnorms, which makes it possible to select highly controlled subsamples$\backslash$nvarying in emotional status. The database is available as a supplement$\backslash$nfor this article at http://brm.psychonomic-journals.org/content/supplemental.
There is a longstanding tradition in psychological research for norming lists of words that are used in experimental studies. The present study extends this practice to graphic imagery by obtaining norming data on 24 simple abstract graphic shapes composed of three straight-line segments. The attributes obtained in the norming procedure were the shapes' familiarity, describability, associability, availability, and potential for word association. Results from rating data indicate significantly different, yet reliable, responses by participants to the various shape configurations. Multidimensional scaling analysis of shape ratings identified two underlying dimensions of perceived differences: the continuity of a shape's linear direction and the consistency or regularity of its interior angles. By contrast, performance in generating word associations for figures appeared to be linguistically driven, with initial responses related to the similarity of shapes to letters of the alphabet. The norms and the computer program used to collect them can be downloaded from www.psychonomic.org/archive. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
Affective stimuli are increasingly used in emotion research. Typically, stimuli are selected from databases providing affective norms. The validity of these norms is a critical factor with regard to the applicability of the stimuli for emotion research. We therefore probed the validity of the Leipzig Affective Norms for German (LANG) by correlating valence and arousal ratings across different sensory modalities. A sample of 120 words was selected from the LANG database, and auditory recordings of these words were obtained from two professional actors. The auditory stimuli were then rated again for valence and arousal. This cross-modal validation approach yielded very high correlations between auditory and visual ratings ({\textgreater}.95). These data confirm the strong validity of the Leipzig Affective Norms for German and encourage their use in emotion research.
Factors that determine unprimed performance in word-stem completion were investigated. Researchers' descriptions of materials used in word-stem completion suggest that normative word frequency, number of alternative completions, and response length are believed to be important. A total of 160 students provided normative responses to each of 914 multiple-completion three-letter word stems, resulting in over 12,000 unique responses. Although the number of alternative responses made correlated well with the number possible, word frequency and length were extremely variable in their relation to response frequency across stems. Creating more precise materials for implicit memory studies appears to require consulting normative response data.
Partly in order to facilitate research on the relation between some standard psychological variables, we gathered normative data on 500 proverbs sampled from theOxford Dictionary of English Proverbs (Wilson, 1970). The scales for which we gathered data are imagery, concreteness, goodness , and familiarity. These norms may be of value to researchers who wish to sample linguistic units larger than the word from a set that contains an extensive number of unfamiliar and familiar items. To illustrate the possible uses to which these data may be put, we presented a causal model of the relation between the four variables mentioned above.
Normative values for word characteristics were obtained from a sample of 12 college-educated, totally congenitally blind subjects on the basis oftheir ratings of 161 nouns on scales of familiar- ity, concreteness, meaningfulness, and imageability. Thedominantmodality ofimagery foreach image-evoking word and the strongest word associate for each item also were recorded. The same data were collected for a group of sighted subjects, both to provide a comparison group for the blind subjects and totest the comparability of sightedsubjects' ratings with existing norms. Rat- ings for sighted subjects correlated strongly with those norms, although the coefficients were slightly higher for ratings of concreteness and imageability than for ratings of familiarity and meaningfulness. Ratings of blind subjects correlated only slightly lower with existing norms for imagery and concreteness, but considerably lower for familiarity and meaningfulness.
Sentence-completion norms for sentences using a multiple production measure are presented. A subset of these items were taken from the Bloom and Fischler (1980) sentence-completion norms in order to compare the Cloze measure with the present multiple production measure. For both measures, the sentence constraint correlated negatively with the number of responses generated across subjects. Although the Cloze measure and the multiple production measure were highly correlated, sentence predictability was higher when the multiple production measure was used. These sentence norms provide an alternative to norms derived using the Cloze procedure.
Prior probabilities of graphemes and conditional probabilities for their pronunciation as specific phonemes are given based on a corpus of 17,310 English words. Phonemes are as given in recent editions ofWebster's New Collegiate Dictionary, with minor revisions; graphemes are defined as letters or letter clusters corresponding to single phonemes. Grapheme-phoneme probabilities were derived from a revised table of frequency of occurrence of phoneme-to-grapheme correspondences generated in a study of spelling regularities (P. R. Hanna, J. S. Hanna, Hodges, {\&} Rudorf, 1966). This quantitative descriptive information provides an index of the strength of particular grapheme-phoneme associations in English. Suggestions are made for the utilization of these probabilities as estimates of spelling/sound predictability in reading research.
Heteronyms are words with 2 different possible pronunciations that are associated with 2 (or more) different meanings. They can be used to investigate psychological mechanisms in reading and other cognitive processes. A corpus of English heteronyms has been collected and is tabulated here. In addition, a corpus of English polyphones is tabulated. These are words with different pronunciations that are not associated with different meanings. (10 ref) (PsycINFO Database Record (c) 2009 APA, all rights reserved)
In a previous article, we presented a systematic computational study of the extraction of semantic representations from the word-word co-occurrence statistics of large text corpora. The conclusion was that semantic vectors of pointwise mutual information values from very small co-occurrence windows, together with a cosine distance measure, consistently resulted in the best representations across a range of psychologically relevant semantic tasks. This article extends that study by investigating the use of three further factors--namely, the application of stop-lists, word stemming, and dimensionality reduction using singular value decomposition (SVD)--that have been used to provide improved performance elsewhere. It also introduces an additional semantic task and explores the advantages of using a much larger corpus. This leads to the discovery and analysis of improved SVD-based methods for generating semantic representations (that provide new state-of-the-art performance on a standard TOEFL task) and the identification and discussion of problems and misleading results that can arise without a full systematic study.
In language-related psychological research, it is often necessary to search for sets of words with certain well-controlled phonological properties. To aid in searches of this kind, this paper presents a matrix of consonant-cluster-free, monosyllabic English words that are classified according to their phonemes. The matrix is of considerable use in the construction of experimental stimuli. Its applications are discussed.
Supplied normative data for the probability of successfully completing 192 single-solution word fragments. Normative data on the familiarity of 80 college students with the solution words were obtained, using ratings, as were estimates of word frequency from existing norms. Regression analyses were performed to predict fragment completion difficulty from familiarity, frequency, and structural characteristics of the fragments. Familiarity, whether or not first and/or last letters appeared in the fragment, and the ratio of letters to missing letters in the fragment were included in the regression equation as significant predictors of difficulty for this fragment set.
Abstract Completion responses were collected for two sets of sentence contexts, which were designed to produce different distributions of probabilities for the primary responses. The subject population consisted of undergraduate college students. For each context, ...$\backslash$n
Ten colors and 10 color words were scaled for meaningfulness by means of Noble's m and a' scale. The most frequent verbal responses to colors and color words were also collected.
Although many recent advances have taken place in corpus-based tools, the techniques used to guide exploration and evaluation of these systems have advanced little. Typically, the plausibility of a semantic space is explored by sampling the nearest neighbors to a target word and evaluating the neighborhood on the basis of the modeler's intuition. Tools for visualization of these large-scale similarity spaces are nearly nonexistent. We present a new open-source tool to plot and visualize semantic spaces, thereby allowing researchers to rapidly explore patterns in visual data that describe the statistical relations between words. Words are visualized as nodes, and word similarities are shown as directed edges of varying strengths. The "Word-2-Word" visualization environment allows for easy manipulation of graph data to test word similarity measures on their own or in comparisons between multiple similarity metrics. The system contains a large library of statistical relationship models, along with an interface to teach them from various language sources. The modularity of the visualization environment allows for quick insertion of new similarity measures so as to compare new corpus-based metrics against the current state of the art. The software is available at www.indiana.edu/{\~{}}semantic/word2word/.
Examined completion responses for 40 3-letter word stems (e.g., ABO) produced by 100 undergraduates. Data include a list of the different words that were written as stem completions, their frequency of occurrence as completions, and their frequency of occurrence in English according to published norms. Analyses revealed 3 primary factors that determined overall performance on a stem-completion test: word frequency, word length, and meanings per word. Usually, however, only 1 of these factors made a significant contribution to performance. It is suggested that results can be used as a database for selecting target words for the construction of completion tests. The entire list of completion responses is appended.
This article presents a new database of 2,654 German nouns rated by a sample of 3,907 subjects on three psycholinguistic attributes: concreteness, valence, and arousal. As a new means of data collection in the field of psycholinguistic research, all ratings were obtained via the Internet, using a tailored Web application. Analysis of the obtained word norms showed good agreement with two existing norm sets. A cluster analysis revealed a plausible set of four classes of nouns: abstract concepts, aversive events, pleasant activities, and physical objects. In an additional application example, we demonstrate the usefulness of the database for creating parallel word lists whose elements match as closely as possible. The complete database is available for free from ftp://ftp.uni-duesseldorf.de/pub/psycho/lahl/WWN. Moreover, the Web application used for data collection is inherently capable of collecting word norms in any language and is going to be released for public use as well.
To conduct experimental investigations into the orthographic processing of Modern Greek, information is needed about the lexical properties known to influence visual word recognition. In this article we introduce GreekLex, a lexical database for Modern Greek, which presents collectively for the first time a series of orthographic measures that can be used for psycholinguistic research. GreekLex consists of 35,304 Modern Greek words ranging in length from 1 to 22 letters, and for each word includes the following statistical information: word length, word-form frequency, lemma frequency, neighborhood density and frequency, transposition neighbors, and addition and deletion neighbors. Furthermore, type and token frequency measures of single letters and bigrams derived from the database are also available. The complete database can be accessed and downloaded freely from www.psychology.nottingham.ac.uk/GreekLex.
This paper provides rating norms for a set of symbols and icons selected from a wide variety of sources. These ratings enable the effects of symbol characteristics on user performance to be systematically investigated. The symbol characteristics that have been quantified are considered to be of central relevance to symbol usability research and include concreteness, complexity, meaningfulness, familiarity, and semantic distance. The interrelationships between each of these dimensions is examined and the importance of using normative ratings for experimental research is discussed.
Pseudowords play an important role in psycholinguistic experiments, either because they are required for performing tasks, such as lexical decision, or because they are the main focus of interest, such as in nonword-reading and nonce-inflection studies. We present a pseudoword generator that improves on current methods. It allows for the generation of written polysyllabic pseudowords that obey a given language's phonotactic constraints. Given a word or nonword template, the algorithm can quickly generate pseudowords that match the template in subsyllabic structure and transition frequencies without having to search through a list with all possible candidates. Currently, the program is available for Dutch, English, German, French, Spanish, Serbian, and Basque, and, with little effort, it can be expanded to other languages.
Dissociations between noun and verb processing are not uncommon after brain injury; yet, precise psycholinguistic comparisons of nouns and verbs are hampered by the underrepresentation of verbs in published semantic word norms and by the absence of contemporary estimates for part-of-speech usage. We report herein imageability ratings and rating response times (RTs) for 1,197 words previously categorized as pure nouns, pure verbs, or words of balanced noun-verb usage on the basis of the Francis and Kucera (1982) norms. Nouns and verbs differed in rated imageability, and there was a stronger correspondence between imageability rating and RT for nouns than for verbs. For all word types, the image-rating-RT function implied that subjects employed an image generation process to assign ratings. We also report a new measure of noun-verb typicality that used the Hyperspace Analog to Language (HAL; Lund {\&} Burgess, 1996) context vectors (derived from a large sample of Usenet text) to compute the mean context distance between each word and all of the pure nouns and pure verbs. For a subset of the items, the resulting HAL noun-verb difference score was compared with part-of-speech usage in a representative sample of the Usenet corpus. It is concluded that this score can be used to estimate the extent to which a given word occurs in typical noun or verb sentence contexts in informal contemporary English discourse. The item statistics given in Appendix B will enable experimenters to select representative examples of nouns and verbs or to compare typical with atypical nouns (or verbs), while holding constant or covarying rated imageability.
This article introduces childLex, an online database of German read by children. childLex is based on a corpus of children's books and comprises 10 million words that were syntactically annotated and lemmatized. childLex reports linguistic norms for lexical, superlexical, and sublexical variables in three different age groups: 6-8 (grades 1-2), 9-10 (grades 3-4), and 11-12 years (grades 5-6). Here, we describe how childLex was collected and analyzed. In addition, we provide information about the distributions of word frequency, word length, and orthographic neighborhood size, as well as their intercorrelations. Finally, we explain how childLex can be accessed using a Web interface.
Age of acquisition (AoA) estimates are provided for 3,460 senses of 1,208 words (i.e., words with multiple meanings e.g., duck). The AoA rating estimates appear to be relatively consistent across participants. The Spearman-Brown split-half reliability coefficient is .95, while the correlations between each participant's ratings and the overall mean ratings yielded correlation coefficients between .325 to .794 with a mean of .69 (SD = .10). These estimates will be of use to those interested in: (a) the influence of AoA on word processing, (b) the influence of AoA on meaning access, (c) the structure of semantic memory, and (d) developmental trends in lexical ambiguity resolution. These AoA estimates can be downloaded from the Psychonomic Society's Web archive of norms, stimuli, and data at www.psychonomic.org/archive.
The psychological community frequently investigates semantic norms of properties produced by native speakers after being presented concept words, and these norms are of great value for a wide variety of psychological experiments. This paper presents a new set of norms that includes a collection of properties from a production experiment for the German and the Italian languages. Stimuli consisted of 50 concrete objects taken from 10 different concept classes. The data comprise annotations of semantic relation types and several statistical measures, which facilitate the comparison of the two target languages.
OBJETIVO: Este estudo comparou os resultados entre crian{\c{c}}as brasileiras e americanas quanto {\`{a}}omea{\c{c}}{\~{a}}o, familiaridade com o conceito representado e complexidade visual de um conjunto de 400 figuras M{\'{E}}TODO: Foram avaliadas 36 crian{\c{c}}as brasileiras (18 meninos) de 5 a 7 anos de idade com caracter{\'{i}}sticas semelhantes {\`{a}}s crian{\c{c}}as americanas. Os procedimentos e medidas empregados no estudo brasileiro foram os mesmos usados para a popula{\c{c}}{\~{a}}o americana permitindo compara{\c{c}}{\~{a}}o direta dos dados das duas amostras atrav{\'{e}}s de correla{\c{c}}{\~{o}}es rho de Spearman e testes t de Student. RESULTADOS: Foram observadas correla{\c{c}}{\~{o}}es positivas significativas para todas as medidas entre as amostras brasileira e americana. A an{\'{a}}lise qualitativa demonstrou que ambos os grupos deram nomes modais que diferem do proposto para 59 figuras. As crian{\c{c}}as brasileiras utilizaram nomes que diferem do proposto para 72 figuras nomeadas corretamente pelas americanas. As americanas nomearam diferentemente do nome modal 26 figuras nomeadas corretamente pelas brasileiras. CONCLUS{\~{A}}O: O conjunto de 400 figuras mostrou-se um instrumento adequado para uso em diferentes culturas. Contudo, {\'{e}} aconselh{\'{a}}vel evitar o uso de figuras que produziram inconsist{\^{e}}ncia de nomea{\c{c}}{\~{a}}o nas popula{\c{c}}{\~{o}}es brasileira e norte-americana em estudos em outras culturas com o mesmo grupo et{\'{a}}rio at{\'{e}} que normas espec{\'{i}}ficas estejam dispon{\'{i}}veis.
Measurements of similarity have typically been obtained through the use of rating, sorting, and perceptual confusion tasks. In the present paper, a new method for measuring similarity is described, in which subjects rearrange items so that their proximity on a computer screen is proportional to their similarity. This method provides very efficient data collection. If a display hasn objects, then, after subjects have rearranged the objects (requiring slightly more thann movements),n(n-1)/2 pairwise similarities can be recorded. As long as the constraints imposed by two-dimensional space are not too different from those intrinsic to psychological similarity, the technique appears to offer an efficient, user-friendly, and intuitive process for measuring psychological similarity.
The present article introduces a Russian-language database of 375 action pictures and associated verbs with normative data. The pictures were normed for name agreement, conceptual familiarity, and subjective visual complexity, and measures of age of acquisition, imageability, and image agreement were collected for the verbs. Values of objective visual complexity, as well as information about verb frequency, length, argument structure, instrumentality, and name relation, are also provided. Correlations between these parameters are presented, along with a comparative analysis of the Russian name agreement norms and those collected in other languages. The full set of pictorial stimuli and the obtained norms may be freely downloaded from http://neuroling.ru/en/db.htm for use in research and for clinical purposes.
We present a database of high-definition (HD) videos for the study of traits inferred from whole-body actions. Twenty-nine actors (19 female) were filmed performing different actions—walking, picking up a box, putting down a box, jumping, sitting down, and standing and acting—while conveying different traits, including four emotions (anger, fear, happiness, sadness), untrustworthiness, and neutral, where no specific trait was conveyed. For the actions conveying the four emotions and untrustworthiness, the actions were filmed multiple times, with the actor conveying the traits with different levels of intensity. In total, we made 2,783 action videos (in both two-dimensional and three-dimensional format), each lasting 7 s with a frame rate of 50 fps. All videos were filmed in a green-screen studio in order to isolate the action information from all contextual detail and to provide a flexible stimulus set for future use. In order to validate the traits conveyed by each action, we asked participants to rate each of the actions corresponding to the trait that the actor portrayed in the two-dimensional videos. To provide a useful database of stimuli of multiple actions conveying multiple traits, each video name contains information on the gender of the actor, the action executed, the trait conveyed, and the rating of its perceived intensity. All videos can be downloaded free at the following address: http://www-users.york.ac.uk/{\~{}}neb506/databases.html. We discuss potential uses for the database in the analysis of the perception of whole-body actions.