1396 norm sets
The temporal characteristics of speech can be captured by examining the distributions of the durations of measurable speech components, namely speech segment durations and pause durations. However, several barriers prevent the easy analysis of pause durations: The first problem is that natural speech is noisy, and although recording contrived speech minimizes this problem, it also discards diagnostic information about cognitive processes inherent in the longer pauses associated with natural speech. The second issue concerns setting the distribution threshold, and consists of the problem of appropriately classifying pause segments as either short pauses reflecting articulation or long pauses reflecting cognitive processing, while minimizing the overall classification error rate. This article describes a fully automated system for determining the locations of speech-pause transitions and estimating the temporal parameters of both speech and pause distributions in natural speech. We use the properties of Gaussian mixture models at several stages of the analysis, in order to identify theoretical components of the data distributions, to classify speech components, to compute durations, and to calculate the relevant statistics.;
Three experiments were conducted to study the effects of enactive imagery (EI) on associative learning. In Experiment I, groups of Ss rated 226 verbs on EI and frequency. In Experiments II and III, Ss learned a 24- and a 16-item list, respectively. The lists consisted of the four possible stimulus-response combinations of high (H) and low (L) EI verb pairs: H-H, H-L, L-H, L-L. In both experiments, EI was found to be a significant factor on the stimulus side, performance being superior when the stimulus was of high EI. In Experiment III, the response EI main effect and the Stimulus by Response EI interaction were also found to be significant. The results indicated that like the imagery evoked by nouns, the EI evoked by verbs facilitates learning.
Although many visual stimulus databases exist, none has data on item similarity levels for multiple items of each kind of stimulus. We present such data for 50 sets of grayscale object photographs. Similarity measures between pictures in each set (e.g., 25 different buttons) were collected using a similarity-sorting method (Goldstone, Behavior Research Methods Instruments {\&} Computers, 26(4):381-386, 1994). A validation experiment used data from 1 picture set and compared responses from standard pairwise measures. This showed close agreement. The similarity-sorting measures were then standardized across picture sets, using pairwise ratings. Finally, the standardized similarity distances were validated in a recognition memory experiment; false alarms increased when targets and foils were more similar. These data will facilitate memory and perception research that needs to make comparisons between stimuli with a range of known target-foil similarities.
Nonverbal vocal expressions, such as laughter, sobbing, and screams, are an important source of emotional information in social interactions. However, the investigation of how we process these vocal cues entered the research agenda only recently. Here, we introduce a new corpus of nonverbal vocalizations, which we recorded and submitted to perceptual and acoustic validation. It consists of 121 sounds expressing four positive emotions (achievement/triumph, amusement, sensual pleasure, and relief) and four negative ones (anger, disgust, fear, and sadness), produced by two female and two male speakers. For perceptual validation, a forced choice task was used (n = 20), and ratings were collected for the eight emotions, valence, arousal, and authenticity (n = 20). We provide these data, detailed for each vocalization, for use by the research community. High recognition accuracy was found for all emotions (86 {\%}, on average), and the sounds were reliably rated as communicating the intended expressions. The vocalizations were measured for acoustic cues related to temporal aspects, intensity, fundamental frequency (f0), and voice quality. These cues alone provide sufficient information to discriminate between emotion categories, as indicated by statistical classification procedures; they are also predictors of listeners' emotion ratings, as indicated by multiple regression analyses. This set of stimuli seems a valuable addition to currently available expression corpora for research on emotion processing. It is suitable for behavioral and neuroscience research and might as well be used in clinical settings for the assessment of neurological and psychiatric patients. The corpus can be downloaded from Supplementary Materials.
In this study, we report normative data by native Persian speakers for concept familiarity, age of acquisition (AoA), imageability, image agreement, name agreement, and visual complexity, as well as values for word frequency, word length, and naming latency for 200 of the colored Snodgrass and Vanderwart (Journal of Experimental Psy- chology: Human Learning and Memory 6:174-215, 1980) pictures created by Rossion and Pourtois (Perception 33:217-236, 2004). Using multiple regression analysis, we found independent effects of name agreement, image agree- ment, word frequency, and AoA on picture naming by native Persian speakers from Iran. We concluded that the psycholinguistic properties identified in studies of picture naming in many other languages also predict timed picture naming in Persian. Normativedatafor theratings and picture-naming latencies for the 200 Persian object nouns are provided as an Excel file in the Supplemental materials.
Subjective frequency estimates for large sample of monosyllabic English words were collected from 574 young adults (undergraduate students) and from a separate group of 1,590 adults of varying ages and educational backgrounds. Estimates from the latter group were collected via the internet. In addition, 90 healthy older adults provided estimates for a random sample of 480 of these words. All groups rated words with respect to the estimated frequency of encounters of each word on a 7-point scale, ranging from never encountered to encountered several times a day. The young and older groups also rated each word with respect to the frequency of encounters in different perceptual domains (e.g., reading, hearing, writing, or speaking). The results of regression analyses indicated that objective log frequency and meaningfulness accounted for most of the variance in subjective frequency estimates, whereas neighborhood size accounted for the least amount of variance in the ratings. The predictive power of log frequency and meaningfulness were dependent on the level of subjective frequency estimates. Meaningfulness was a better predictor of subjective frequency for uncommon words, whereas log frequency was a better predictor of subjective frequency for common words. Our discussion focuses on the utility of subjective frequency estimates compared with other estimates of familiarity. The raw subjective frequency data for all words are available at http://www.artsci.wustl.edu/dbalota/labpub.html.
Complete normative data are presented for responses to 56 verbal categories by students from the Universities of Maryland (W=270) and Illinois (JV= 172). All 43 categories from the Connecticut norms are included, and complete data are presented for all responses given by each S to each category label within 30 sec. Additional data include (a) number of times each response was given first and mean rank of each response, (6) correlations between the various measures and between the Maryland and Illinois samples for each category, and (c) "category potency" measures and ratings for each category.
We report object-naming and object recognition times collected from Russian native speakers for the colorized version of the Snodgrass and Vanderwart (Journal of Experimental Psychology: Human Learning and Memory 6:174-215, 1980) pictures (Rossion {\&} Pourtois, Perception 33:217-236, 2004). New norms for image variability, body-object interaction [BOI], and subjective frequency collected in Russian, as well as new name agreement scores for the colorized pictures in French, are also reported. In both object-naming and object comprehension times, the name agreement, image agreement, and age-of-acquisition variables made significant independent contributions. Objective word frequency was reliable in object-naming latencies only. The variables of image variability, BOI, and subjective frequency were not significant in either object naming or object comprehension. Finally, imageability was reliable in both tasks. The new norms and object-naming and object recognition times are provided as supplemental materials.
We report psycholinguistic norms for 305 French idiomatic expressions (Study 1). For each of the idiomatic expressions, the following variables are reported: knowledge, predictability, literality, compositionality, subjective and objective frequency, familiarity, age of acquisition (AoA), and length. In addition, we have collected comprehension times for each idiom (Study 2). The psycholinguistic relevance of the collected norms is explained, and different analyses (descriptive statistics, correlation and multiple regression analyses) performed on the norms are reported and discussed. The entire set of norms and reading times are provided as supplemental material.
We present a new database of lexical decision times for English words and nonwords, for which two groups of British participants each responded to 14,365 monosyllabic and disyllabic words and the same number of nonwords for a total duration of 16 h (divided over multiple sessions). This database, called the British Lexicon Project (BLP), fills an important gap between the Dutch Lexicon Project (DLP; Keuleers, Diependaele, {\&} Brysbaert, Frontiers in Language Sciences. Psychology, 1, 174, 2010) and the English Lexicon Project (ELP; Balota et al., 2007), because it applies the repeated measures design of the DLP to the English language. The high correlation between the BLP and ELP data indicates that a high percentage of variance in lexical decision data sets is systematic variance, rather than noise, and that the results of megastudies are rather robust with respect to the selection and presentation of the stimuli. Because of its design, the BLP makes the same analyses possible as the DLP, offering researchers with a new interesting data set of word-processing times for mixed effects analyses and mathematical modeling. The BLP data are available at http://crr.ugent.be/blp and as Electronic Supplementary Materials.
Many studies have addressed the issue of whether age of acquisition and/ or frequency affect particular lexical tasks. Methods typically employed in such studies are based on the general linear model (e.g., ANOVA or multiple regression). These methods assume manipulated independent variables whereas the usual approach of investigating age-of-acquisition and frequency effects uses estimated norms of word properties. This failure to truly manipulate variables violate the assumptions of the analyses. A simulation is provided that demonstrates how this violation can lead to erroneous conclusions of effects when none are present. Recommendations are made for a more correlational approach to analysis using structural equation modelling techniques. It is also discussed how this use of estimates of lexical data is problematic for determining effects throughout psycholinguistic research. {\textcopyright} 2006 Psychology Press Ltd.
Studies of lexical processing have relied heavily on adult ratings of word learning age or age of acquisition, which have been shown to be strongly predictive of processing speed. This study reports a set of objective norms derived in a large-scale study of British children's naming of 297 pictured objects (including 232 from the Snodgrass {\&} Vanderwart, 1980, set). In addition, data were obtained on measures of rated age of acquisition, rated frequency, imageability, object familiarity, picture-name agreement, and name agreement. We discuss the relationship between the objective measure and adult ratings of word learning age. Objective measures should be used when available, but where not, our data suggest that adult ratings provide a reliable and valid measure of real word learning age.
In Exp I, 328 adjectives were presented to 324 undergraduates for rating of imagery (I), ease of definition (ED), and animateness (A). The normative value of these indices was tabulated for each adjective. A correlational analysis of these measures and Ku{\v{c}}era-Francis frequency (KF) is also presented. To demonstrate the usefulness of these rating scales, Exp II requested 90 undergraduates to free recall a list of 50 adjectives after they completed an incidental learning task of rating these adjectives for I, ED, or A. Ss recalled more high I than low I adjectives and more difficult- than easy-to-define adjectives regardless of which incidental rating task they performed. Neither the degree of animateness nor the KF value of the adjectives influenced the percentage of recall.
Focuses on the statistical relationship established between high frequency, small variety and shortness in length of words. Analysis of written language; Influence of age differences with the relationship between variety and frequency of occurrence of words; Growth of available vocabulary.
A list of gender-related and gender-neutral words for use in testing gender stereotyping and memory was created and evaluated. Words were rated by samples of undergraduates at universities located in the northeast, southeast, and south-central United States. A substantial list of masculine, feminine, and gender-neutral words was identified. These lists allow researchers to construct large lists of gender-associated words while being able to control for extraneous variables, such as word frequency and word length. In addition, the high reliability across the samples suggests that gender ratings are a fairly stable phenomenon. Applications for this list are discussed. The word lists presented in Tables 1-3 and the raw data analyzed in this article may be downloaded from www.psychonomic.org/archive/.
This paper presents an analysis of the distribution of phonological similarity relations among monosyllabic spoken words in English. It differs from classical analyses of phonological neighborhood density (e.g., Luce {\&} Pisoni, 1998) by assuming that not all phonological neighbors are equal. Rather, it is assumed that the phonological lexicon has psycholinguistic structure. Accordingly, in addition to considering the number of phonological neighbors for any given word, it becomes important to consider the nature of these neighbors. If one type of neighbor is more dominant, neighborhood density effects may reflect levels of segmental representation other than the phoneme, particularly prior to literacy. Statistical analyses of the nature of phonological neighborhoods in terms of rime neighbors (e.g., hat/cat), consonant neighbors (e.g., hat/hit), and lead neighbors (e.g., hat/ham) were thus performed for all monosyllabic words in the Celex corpus (4,086 words). Our results show that most phonological neighbors are rime neighbors (e.g., hat/cat) in English. Similar patterns were found when a corpus of words for which age-of-acquisition ratings were available was analyzed. The resultant database can be used as a tool for controlling and selecting stimuli when the role of lexical neighborhoods in phonological development and speech processing is examined.
The Nelson and Narens (Journal of Verbal Learning and Verbal Behavior 19:338-368, 1980) general knowledge norms have been valuable to researchers in many fields. However, much has changed over the 32 years since the 1980 norms. For example, in 1980, most people knew the answer to the question "What is the name of the Lone Ranger's Indian sidekick?" (answer: Tonto), whereas in 2012, few people know this answer. Thus, we updated the 1980 norms and expanded them by providing new measures. In particular, we report two new metacognitive measures (confidence judgments and peer judgments) and provide a detailed report of commission errors. Each of these measures will be valuable to researchers, and together they are likely to facilitate future research in a number of fields, such as research investigating memory illusions, metamemory processes, and error correction. The presence of substantial generational shifts from 1980 to 2012 necessitates the use of updated norms.
Age of acquisition (AoA) ratings made on a 1–7 scale for 3,000 monosyllabic words were obtained from 32 participants across four blocks of 750 trials (two blocks of 750 trials were completed in each of 2 days). These results, as well as those of the regression analyses and reliability and validity measures that were originally reported in Cortese and Khanna (2007), are summarized here. Here, we also report high interblock correlations across items, indicating that participants were consistent in their ratings across blocks. The norms for the 3,000 words are important for researchers interested in word processing and may be downloaded from the Psycho- nomic Society's Norms, Stimuli, and Data archive at www.psychonomic.org/archive.
The use of rhyme in learning/memory and cognitive studies is extensive. However, there are very few normative studies for words that rhyme. The current study rectified this problem by collecting rhyme norms for 477 words from 545 subjects. Groups of subjects were given 40 words in serial order and requested, for each word, to generate as many rhymes as possible within a 30-sec interval. The data include several rhyme measures as well as measures of other word attributes that were taken from other sources. In addition, the rhyme responses to the target words were given along with their Thomdike and Lorge (1944) and Ku{\v{c}}era and Francis (1967) normative frequencies. Finally, the data were used to investigate various relationships including the spew hypothesis and the accessibility of rhyme sets.
In this article, we introduce ESCOLEX, the first European Portuguese children's lexical database with grade-level-adjusted word frequency statistics. Computed from a 3.2-million-word corpus, ESCOLEX provides 48,381 word forms extracted from 171 elementary and middle school textbooks for 6- to 11-year-old children attending the first six grades in the Portuguese educational system. Like other children's grade-level databases (e.g., Carroll, Davies, {\&} Richman, 1971; Corral, Ferrero, {\&} Goikoetxea, Behavior Research Methods, 41, 1009–1017, 2009; L{\'{e}}t{\'{e}}, Sprenger-Charolles, {\&} Col{\'{e}}, Behavior Research Methods, Instruments, {\&} Computers, 36, 156–166, 2004; Zeno, Ivens, Millard, Duvvuri, 1995), ESCOLEX provides four frequency indices for each grade: overall word frequency (F), index of dispersion across the selected textbooks (D), estimated frequency per million words (U), and standard frequency index (SFI). It also provides a new measure, contextual diversity (CD). In addition, the number of letters in the word and its part(s) of speech, number of syllables, syllable structure, and adult frequencies taken from P-PAL (a European Portuguese corpus-based lexical database; Soares, Comesa{\~{n}}a, Iriarte, Almeida, Sim{\~{o}}es, Costa, {\ldots}, Machado, 2010; Soares, Iriarte, Almeida, Sim{\~{o}}es, Costa, Fran{\c{c}}a, {\ldots}, Comesa{\~{n}}a, in press) are provided. ESCOLEX will be a useful tool both for researchers interested in language processing and development and for professionals in need of verbal materials adjusted to children's developmental stages. ESCOLEX can be downloaded along with this article or from http://p-pal.di.uminho.pt/about/databases.
In this study, we present the normative values of the adaptation of the International Affective Digitized Sounds (IADS-2; Bradley {\&} Lang, 2007a) for European Portuguese (EP). The IADS-2 is a standardized database of 167 naturally occurring sounds that is widely used in the study of emotions. The sounds were rated by 300 college students who were native speakers of EP, in the three affective dimensions of valence, arousal, and dominance, by using the Self-Assessment Manikin (SAM). The aims of this adaptation were threefold: (1) to provide researchers with standardized and normatively rated affective sounds to be used with an EP population; (2) to investigate sex and cultural differences in the ratings of affective dimensions of auditory stimuli between EP and the American (Bradley {\&} Lang, 2007a) and Spanish (Fern{\'{a}}ndez-Abascal et al., Psicothema 20:104-113 2008; Redondo, Fraga, Padr{\'{o}}n, {\&} Pi{\~{n}}eiro, Behavior Research Methods 40:784-790 2008) standardizations; and (3) to promote research on auditory affective processing in Portugal. Our results indicated that the IADS-2 is a valid and useful database of digitized sounds for the study of emotions in a Portuguese context, allowing for comparisons of its results with those of other international studies that have used the same database for stimulus selection. The normative values of the EP adaptation of the IADS-2 database can be downloaded along with the online version of this article.
In this article we present a standardized set of 260 pictures for use in experiments investigating differences and similarities in the processing of pictures and words. The pictures are black-and-white line drawings executed according to a set of rules that provide consistency of pictorial representation. The pictures have been standardized on four variables of central relevance to memory and cognitive processing: name agreement, image agreement, familiarity, and visual complexity. The intercorrelations among the four measures were low, suggesting that they are indices of different attributes of the pictures. The concepts were selected to provide exemplars from several widely studied semantic categories. Sources of naming variance, and mean familiarity and complexity of the exemplars, differed significantly across the set of categories investigated. The potential significance of each of the normative variables to a number of semantic and episodic memory tasks is discussed.
Ratings are presented for 650 stimuli from word-association lists on each of five scales: good—bad, pleasant—unpleasant, emotional—neutral, concrete—abstract, and easy to associate to—difficult to associate to. The ratings are shown to be highly reliable, and to agree well with previously collected norms of a similar character. The intercorrelations of the five scales with one another and with word frequency are reported.
A count of the relative incidence of letters in 431 pleasant (P) words and 702 unpleasant (U) words revealed that some letters tend to occur more frequently in the initial position of P words and other letters more frequently in the initial position of U words. A task requiring Ss to guess for each letter whether it occurred in a P word or in a U word showed that people are able to approximate these objective probabilities of initial-letter occurrence. These findings can explain how it is possible to identify the probable affective meaning of a word seen in a tachistoscope prior to its complete recognition. Faster recognition of P words in tachistoscopic experiments was accounted for in terms of response probability. {\textcopyright} 1969 Academic Press Inc. All rights reserved.
Perceptual cues mediating recognition of isolated lowercase letters have been investigated in two conditions of marginal reading:from a long distance and in eccentric vision. A high incidence of confusions indicates that observers readily use available cues for arriving at letter responses. Analysis of the confusions leads to perceptual similarities and to common properties that have possibly served as perceptual cues. Dominating similarities are /h k b/, /tilfr/,/eoc/,/aszxe/,/vw/and/gq/. Properties with high cue values are: 1. (1) vertically ascending and descending parts, 2. (2) slenderness, 3. (3) outer vertical and outer oblique parts, and also outer gaps. Inner parts are weak cues at best. Bias effects occur towards letters that occur frequently in the printed language, but they are restricted to confusions. Perceptual cues prevail over bias effects. In eccentric vision, recognition is limited by more factors than just a low visual acuity. {\textcopyright} 1971.
INFORMATION regarding the sequential dependencies among letters in the English language is of interest not only in connexion with linguistics and cryptography, but also because of its value for communication theory and for the psychology of language1,2. ? 1960 Nature Publishing Group.
Throughout the last decades, numerous picture data sets have been developed, such as the Snodgrass and Vanderwart (1980) set, and have been normalized for variables such as name and familiarity; however, due to cultural and linguistic differences, norms can vary from one country to another. The effect due specifically to culture has already been demonstrated by comparing samples from different countries where the same language is spoken. On the other hand, it is still not clear how differences between languages may affect norms. The present study explores this issue by collecting and comparing norms on names and many other features from French Canadian speakers and English Canadian speakers living in Montreal, who thus live in similar cultural environments. Norms were collected for the photos of objects from the Bank of Standardized Stimuli (BOSS) by asking participants to name the objects, to categorize them, and to rate their familiarity, visual complexity, object agreement, viewpoint agreement, and manipulability. Names and ratings from the French speakers are available in Appendix A, available in the supplemental materials. The results show that most of the norms are comparable across linguistic groups and also that the ratings given are correlated across linguistic groups. The only significant group differences were found in viewpoint agreement and visual complexity. Overall, there was good concordance between the norms collected from French and English native speakers living in the same cultural setting.
Ratings of the association values on a five-point scale by 95 Ss are reported for the 101 numbers between 0 and 100. The results reveal sizable and relatively consistent differences between numbers which are shown to be related to learning difficulty, as well as a high degree of individual variability. {\textcopyright} 1962.
We present new Spanish norms for object familiarity and rated age of acquisition for 140 pictures taken from Snodgrass and Vanderwart (1980), together with data on visual complexity, image agreement, name agreement, word length (in syllables and phonemes), and five measures of word frequency. The pictures were presented to a group of 64 Spanish subjects, and oral naming latencies were recorded. In a multiple regression analysis, age of acquisition, object familiarity, name agreement, word frequency, and word length made significant independent contributions to predicting naming latency.
Compared sentence completion responses across 130 adults, aged 18-56 yrs old, for 198 highly constrained sentence contexts designed to elicit the best completion in the vast majority of Subjects. For each context, completions and their respective frequency of occurrence are provided. Subjects of all ages produced highly similar terminal words. Results indicate that greater SES and higher levels of education were mildly associated with a greater probability of producing a best completion response. Although increasing age correlated with greater probability of producing a best completion, this very weak association would not preclude use of these stimuli with a wide age range. (PsycINFO Database Record (c) 2000 APA, all rights reserved)
The Oxford English Dictionary is the standard reference work for determining the earliest known instance of the occurrence of a word (its date of entry). Partly in order to facilitate research on the relation between the date of entry and other psychological variables, we gathered normative data on 1,046 words sampled from the Oxford English Dictionary. We present data for scales that measure the imagery, concreteness, goodness, and familiarity values for words. These norms may also be of use to researchers who are not explicitly concerned with words' date of entry, but who wish to sample words from a set that contains a large number of unfamiliar as well as familiar words.
The general aim of this study is to validate the cognitive relevance of the geometric model used in the semantic atlases (SA). With this goal in mind, we compare the results obtained by the automatic contexonym organizing model (ACOM)--an SA-derived model for word sense representation based on contextual links--with human subjects' responses on a word association task. We begin by positioning the geometric paradigm with respect to the hierarchical paradigm (WordNet) and the vector paradigm (latent semantic analysis [LSA] and the hyperspace analogue to language model). Then we compare ACOM's responses with Hirsh and Tree's (2001) word association norms based on the responses of two groups of subjects. The results showed that words associated by 50{\%} or more of the Hirsh and Tree subjects were also proposed by ACOM (e.g., 71{\%} of the words in the norms were also given by ACOM). Finally, we compare ACOM and LSA on the basis of the same association norms. The results indicate better performance for the geometric model.
WordNet, an electronic dictionary (or lexical database), is a valuable resource for computational and cognitive scientists. Recent work on the computing of semantic distances among nodes (synsets) in WordNet has made it possible to build a large database of semantic distances for use in selecting word pairs for psychological research. The database now contains nearly 50,000 pairs of words that have values for semantic distance, associative strength, and similarity based on co-occurrence. Semantic distance was found to correlate weakly with these other measures but to correlate more strongly with another measure of semantic relatedness, featural similarity. Hierarchical clustering analysis suggested that the knowledge structure underlying semantic distance is similar in gross form to that underlying featural similarity. In experiments in which semantic similarity ratings were used, human participants were able to discriminate semantic distance. Thus, semantic distance as derived from WordNet appears distinct from other measures of word pair relatedness and is psychologically functional. This database may be downloaded from www.psychonomic.org/archive/.
Six previous studies of the variables affecting anagram solution are re-examined for the evidence that number of syllables contributes to solution difficulty. It was shown that the number of syllables in a solution word was confounded with imagery for one study and with diagram frequency for another. More importantly it was shown that the number of syllables has a large effect on anagram solution difficulty in the re-analysis of the results from the other four studies. In these studies, the number of syllables was either more important than the principal variable examined in the experiment or the second most important variable. Overall the effect size for the number of syllables was large, d = 1.14. The results are discussed in the light of other research and it is suggested that anagram solution may have more in common with other word identification and reading processes than has been previously thought.
Cognitive psychology is finding increasing use for the word fragment completion test, in which words have to be completed from a subset of their letters (e.g., horizon from --r-z--). Researchers often try to restrict their fragments to those that can be completed with only one word, but this is difficult to do and probably never has been achieved. To help resolve this problem, a list is provided of 1,086 three- to eight-letter words, each of which is uniquely specified by a two-letter fragment, where uniqueness is defined on the basis of two sizable word collections.
The processes involved in past tense verb generation have been central to models of inflectional morphology. However, the empirical support for such models has often been based on studies of accuracy in past tense verb formation on a relatively small set of items. We present the first large-scale study of past tense inflection (the Past Tense Inflection Project, or PTIP) that affords response time, accuracy, and error analyses in the generation of the past tense form from the present tense form for over 2,000 verbs. In addition to standard lexical variables (such as word frequency, length, and orthographic and phonological neighborhood), we have also developed new measures of past tense neighborhood consistency and verb imageability for these stimuli, and via regression analyses we demonstrate the utility of these new measures in predicting past tense verb generation. The PTIP can be used to further evaluate existing models, to provide well controlled stimuli for new studies, and to uncover novel theoretical principles in past tense morphology.
In prior experiments S-generated associative devices or natural language mediators (NLMs) linking pairs of items have been shown to facilitate acquisition of paired associates. Since Ss are questioned about NLMs after learning, such repos/rts may be a result of the questioning. To obtain an a priori estimate of NLM probability, several hundred pairs, each composed of CVCs of about equal association value (AV), were shown for 15 sec. while student Ss wrote down any NLM they could generate which linked both the stumulus and response. The AV level was varied between pairs. The proportion of Ss able to generate an NLM is the associability value (AS). As expected, AS and AV are correlated although AS varies considerably among pairs composed of items about equal in AV. Experiments run after the AS scale was obtained demonstrated that AS is valuable as a predictor of learning rate. AS values were highly correlated with the frequency of NLMs in postexperiment reports. It is concluded that the AS measure represents a valuable addition to our understanding of the complexity of verbal learning.
We report normative data collected from Mainland Chinese speakers for 232 objects taken from Snodgrass and Vanderwart (1980). These data include adult ratings of concept familiarity, age of acquisition (AoA), printedword frequency, and word length (in syllables), as well as measures of rated visual complexity, image agreement, and name agreement. We then examined timed picture naming of these objects with native Chinese speakers in Beijing in two experiments using line drawings and colored pictures. In both experiments, the variables name agreement, rated concept familiarity, and AoA made significant independent contributions to naming latency in multiple regression analyses. We observed a correlation ofr=.85 between naming latency with line drawings and colored pictures and a reduced effect of image agreement on naming when colored pictures were presented. We discuss the implications of our findings for the study of lexical processing in Chinese. Normative data for 232 Chinese nouns may be downloaded from www.psychonomic.org/archive
Phonotactic probability refers to the frequency with which phonological segments and sequences of phonological segments occur in words in a given language. We describe one method of estimating phonotactic probabilities based on words in American English. These estimates of phonotactic probability have been used in a number of previous studies and are now being made available to other researchers via a Web-based interface. Instructions for using the interface, as well as details regarding how the measures were derived, are provided in the present article. The Phonotactic Probability Calculator can be accessed at http://www.people.ku.edu/{\~{}}mvitevit/PhonoProbHome.html.
The present article provides Spanish norms for name agreement, printed word frequency, word compound frequency, familiarity, imageability, visual complexity, age of acquisition, and word length (measured by syllables and phonemes) for 100 line drawings of actions taken from Druks and Masterson (2000). In addition, through a naming-time experiment carried out with a group of 54 Spanish students in a pool of 63 of these line drawings, we determined the best predictors of naming actions. In the multiple regression analysis, age of acquisition and name agreement emerged as the most important determinants of action-naming reaction time.
This study introduces the Tool for the Automatic Analysis of Cohesion (TAACO), a freely available text analysis tool that is easy to use, works on most operating systems (Windows, Mac, and Linux), is housed on a user's hard drive (rather than having an Internet interface), allows for the batch processing of text files, and incorporates over 150 classic and recently developed indices related to text cohesion. The study validates TAACO by investigating how its indices related to local, global, and overall text cohesion can predict expert judgments of text coherence and essay quality. The findings of this study provide predictive validation of TAACO and support the notion that expert judgments of text coherence and quality are either negatively correlated or not predicted by local and overall text cohesion indices, but are positively predicted by global indices of cohesion. Combined, these findings provide supporting evidence that coherence for expert raters is a property of global cohesion and not of local cohesion, and that expert ratings of text quality are positively related to global cohesion.
In emotional research, efficient designs often rely on successful emotion induction. For visual stimulation, the only reliable database available so far is the International Affective Picture System (IAPS). However, extensive use of these stimuli lowers the impact of the images by increasing the knowledge that participants have of them. Moreover, the limited number of pictures for specific themes in the IAPS database is a concern for studies centered on a specific emotion thematic and for designs requiring a lot of trials from the same kind (e.g., EEG recordings). Thus, in the present article, we present a new database of 730 pictures, the Geneva Affective PicturE Database, which was created to increase the availability of visual emotion stimuli. Four specific negative contents were chosen: spiders, snakes, and scenes that induce emotions related to the violation of moral and legal norms (human rights violation or animal mistreatment). Positive and neutral pictures were also included: Positive pictures represent mainly human and animal babies as well as nature sceneries, whereas neutral pictures mainly depict inanimate objects. The pictures were rated according to valence, arousal, and the congruence of the represented scene with internal (moral) and external (legal) norms. The constitution of the database and the results of the picture ratings are presented.
Lexvo.org brings information about languages, words, and other linguistic entities to the Web of Linked Data. It defines URIs for terms, languages, scripts, and characters, which are not only highly interconnected but also linked to a variety of resources on the Web. Additionally, new datasets are being published to contribute to the emerging Linked Data Cloud of Language-Related information.
This paper describes the Atlante Sintattico d'Italia, Syntactic Atlas of Italy (ASIt) linguistic linked dataset. ASIt is a scientific project aiming to account for minimally different variants within a sample of closely related languages; it is part of the Edisyn network, the goal of which is to establish a European network of researchers in the area of language syntax that use similar standards with respect to methodology of data collection, data storage and annotation, data retrieval and cartography. In this context, ASIt is defined as a curated database which builds on dialectal data gathered during a twenty-year-long survey investigating the distribution of several grammatical phenomena across the dialects of Italy. Both the ASIt linguistic linked dataset and the Resource Description Framework Schema (RDF/S) on which it is based are publicly available and released with a Creative Commons license (CC BY-NC-SA 3.0). We report the characteristics of the data exposed by ASIt, the statistics about the evolution of the data in the last two years, and the possible usages of the dataset, such as the generation of linguistic maps. {\textcopyright} 2012 - IOS Press and the authors. All rights reserved.
As human activity and interaction increasingly take place online, the digital residues of these activities provide a valuable window into a range of psychological and social processes. A great deal of progress has been made toward utilizing these opportunities; however, the complexity of managing and analyzing the quantities of data currently available has limited both the types of analysis used and the number of researchers able to make use of these data. Although fields such as computer science have developed a range of techniques and methods for handling these difficulties, making use of those tools has often required specialized knowledge and programming experience. The Text Analysis, Crawling, and Interpretation Tool (TACIT) is designed to bridge this gap by providing an intuitive tool and interface for making use of state-of-the-art methods in text analysis and large-scale data management. Furthermore, TACIT is implemented as an open, extensible, plugin-driven architecture, which will allow other researchers to extend and expand these capabilities as new methods become available.
This paper presents a rule-based approach for generating a large phonetic database for Romanian. The knowledge base is developed by means of the GRAALAN (Grammar Abstract Language) system. By inspecting dictionaries and corpora, we generate a phonetic database over 100,000 lemmas. Our database has a high degree of accuracy ensured by our rule-based method applied for generating phonetic transcriptions.
Despite their relatively low sampling factor, the freely available, randomly sampled status streams of Twitter are very useful sources of geographically embedded social network data. To statistically analyze the information Twitter provides via these streams, we have collected a year's worth of data and built a multi-terabyte relational database from it. The database is designed for fast data loading and to support a wide range of studies focusing on the statistics and geographic features of social networks, as well as on the linguistic analysis of tweets. In this paper we present the method of data collection, the database design, the data loading procedure and special treatment of geo-tagged and multi-lingual data. We also provide some SQL recipes for computing network statistics.
Word lists are most commonly used in the investigation of human memory. To prevent transfer effects, repeated measures of memory for words require multiple lists of different words. Yet, the psycholinguistic properties of all word lists employed should match as closely as possible to avoid confounding with the independent variable(s) in question. Although comprehensive databases for word norms exist, to our knowledge no tool is available that automates the creation of such equivalent word lists. Instead, matching different lists is often accomplished prima facie. We have therefore developed a Windows program called EQUIWORD that completely automates the creation of word lists that are truly parallel with respect to a wide range of attributes. EQUIWORD takes psycholinguistic databases of different formats as input and computes several coefficients of distance for every possible word pairing. Program output consists of a list of all word pairs sorted according to their distance. On that basis, creating equivalent word lists is simply done by selecting the pairs with the lowest distance coefficients.
We present a database of 858 German words from the semantic fields of authority and community, which represent core dimensions of human sociality. The words were selected on the basis of co-occurrence profiles of representative keywords for these semantic fields. All words were rated along five dimensions, each measured by a bipolar semantic-differential scale: Besides the classic dimensions of affective meaning (valence, arousal, and potency), we collected ratings of authority and community with newly developed scales. The results from cluster, correlational, and multiple regression analyses on the rating data suggest a robust negativity bias for authority valuation among German raters recruited via university mailing lists, whereas community ratings appear to be rather unrelated to the well-established affective dimensions. Furthermore, our data involve a strong overall negative correlation-rather than the classical U-shaped distribution-between valence and arousal for socially relevant concepts. Our database provides a valuable resource for research questions at the intersection of cognitive neuroscience and social psychology. It can be downloaded as supplemental materials with this article.
Although many facial and vocal databases are available for research, very few of them have controlled the range of attractiveness of the stimuli that they offer. To fill this gap, we created the GEneva Faces and Voices (GEFAV) database, providing standardized faces (static and dynamic neutral, smiling) and voices (speaking sentences, vowels) of young European adults. A total of 61 women and 50 men 18-35 years old agreed to be part of the GEFAV stimuli, and two rating studies involving 285 participants provided evaluations of the facial and vocal samples. The final set of stimuli was satisfactory in terms of attractiveness range (wide and rather symmetrical distribution over the attractiveness continuum) and the reliability of the ratings (high consistency between the two rating studies, high interrater agreement in the final rating study). Moreover, the database showed an adequate validity, since a series of findings described by earlier research on human attractiveness were confirmed-namely, that facial and vocal attractiveness are predicted by femininity and health in women, and by masculinity, dominance, and trustworthiness in men. In future studies, the GEFAV stimuli may be used intact or transformed, individually or in multimodal combinations, to investigate a wide range of mechanisms, such as the behavioral, neuropsychological, and neurophysiological processes involved in social cognition.