1396 norm sets
Research into similarities between music and language processing is currently experiencing a strong renewed interest. Recent methodological advances have led to neuroimaging studies presenting striking similarities between neural patterns associated with the processing of music and language--notably, in the study of participants' responses to elements that are incongruous with their musical or linguistic context. Responding to a call for greater systematicity by leading researchers in the field of music and language psychology, this article describes the creation, selection, and validation of a set of auditory stimuli in which both congruence and resolution were manipulated in equivalent ways across harmony, rhythm, semantics, and syntax. Three conditions were created by changing the contexts preceding and following musical and linguistic incongruities originally used for effect by authors and composers: Stimuli in the incongruous-resolved condition reproduced the original incongruity and resolution into the same context; stimuli in the incongruous-unresolved condition reproduced the incongruity but continued postincongruity with a new context dictated by the incongruity; and stimuli in the congruous condition presented the same element of interest, but the entire context was adapted to match it so that it was no longer incongruous. The manipulations described in this article rendered unrecognizable the original incongruities from which the stimuli were adapted, while maintaining ecological validity. The norming procedure and validation study resulted in a significant increase in perceived oddity from congruous to incongruous-resolved and from incongruous-resolved to incongruous-unresolved in all four components of music and language, making this set of stimuli a theoretically grounded and empirically validated resource for this growing area of research.
The discrete emotion theory proposes that affective experiences can be reduced to a limited set of universal "basic" emotions, most commonly identified as happiness, sadness, anger, fear, and disgust. Here we present norms for 10,491 Spanish words for those five discrete emotions collected from a total of 2,010 native speakers, making it the largest set of norms for discrete emotions in any language to date. When used in conjunction with the norms from Hinojosa, Mart{\'{i}}nez-Garc{\'{i}}a et al. (Behavior Research Methods, 48, 272-284, 2016) and Ferr{\'{e}}, Guasch, Mart{\'{i}}nez-Garc{\'{i}}a, Fraga, {\&} Hinojosa (Behavior Research Methods, 49, 1082-1094, 2017), researchers now have access to ratings of discrete emotions for 13,633 Spanish words. Our norms show a high degree of inter-rater reliability and correlate highly with those from Ferr{\'{e}} et al. (2017). Our exploration of the relationship between the five discrete emotions and relevant lexical and emotional variables confirmed findings of previous studies conducted with smaller datasets. The availability of such large set of norms will greatly facilitate the study of emotion, language and related fields. The norms are available as supplementary materials to this article.
In everyday social interactions, people's facial expressions sometimes reflect genuine emotion (e.g., anger in response to a misbehaving child) and sometimes do not (e.g., smiling for a school photo). There is increasing theoretical interest in this distinction, but little is known about perceived emotion genuineness for existing facial expression databases. We present a new method for rating perceived genuineness using a neutral-midpoint scale (-7 = completely fake; 0 = don't know; +7 = completely genuine) that, unlike previous methods, provides data on both relative and absolute perceptions. Normative ratings from typically developing adults for five emotions (anger, disgust, fear, sadness, and happiness) provide three key contributions. First, the widely used Pictures of Facial Affect (PoFA; i.e., "the Ekman faces") and the Radboud Faces Database (RaFD) are typically perceived as not showing genuine emotion. Also, in the only published set for which the actual emotional states of the displayers are known (via self-report; the McLellan faces), percepts of emotion genuineness often do not match actual emotion genuineness. Second, we provide genuine/fake norms for 558 faces from several sources (PoFA, RaFD, KDEF, Gur, FacePlace, McLellan, News media), including a list of 143 stimuli that are event-elicited (rather than posed) and, congruently, perceived as reflecting genuine emotion. Third, using the norms we develop sets of perceived-as-genuine (from event-elicited sources) and perceived-as-fake (from posed sources) stimuli, matched on sex, viewpoint, eye-gaze direction, and rated intensity. We also outline the many types of research questions that these norms and stimulus sets could be used to answer.
Equal numbers of male and female participants judged which of seven facial expressions (anger, disgust, fear, happiness, neutrality, sadness, and surprise) were displayed by a set of 336 faces, and we measured both accuracy and response times. In addition, the participants rated how well the expression was displayed (i.e., the intensity of the expression). These three measures are reported for each face. Sex of the rater did not interact with any of the three measures. However, analyses revealed that some expressions were recognized more accurately in female than in male faces. The full set of these norms may be downloaded fromwww.psychonomic.org/archive/.
We present a set of stimuli representing human actions under point-light conditions, as seen from different viewpoints. The set contains 22 fairly short, well-delineated, and visually “loopable” actions. For each action, we provide movie files from five different viewpoints as well as a text file with the three spatial coordinates of the point lights, allowing researchers to construct customized versions. The full set of stimuli may be downloaded fromwww.psychonomic.org/archive/.
Describes software that provides a set of pictorial stimuli prepared for computer presentation circumventing the time-consuming nature of individually drawing pictures on the computer. Access of the stimuli from Microsoft Basic, hardware and software requirements, and availability of the software are discussed.
The present study provides normative measures for a new stimulus set of images consisting of 225 everyday objects, each depicted both as a photograph and a matched clipart image generated directly from the photograph (450 images total). The clipart images preserve the same scale, shape, orientation, and general color features as the corresponding photographs. Various norms (modal name and verb agreement measures, picture-name agreement, familiarity, visual complexity, and image agreement) were collected separately for each image type and in two different contexts: online (using Mechanical Turk) and in the laboratory. We discuss similarities and differences in the normative measures according to both image type and experimental context. The full set of norms is provided in the supplemental materials.
Picture databases are commonly used in experimental work on various aspects of emotion processing. However, existing standardized facial databases, typically used to explore emotion recognition, can be augmented with more contextual information for studying emotion and social perception. Moreover, the perception of social engagement, i.e., the degree of interaction or engagement inferred between the people in target pictures, has not been measured. In this paper, we describe the development of a database comprising 203 black-and-white line drawings depicting people within various situational contexts, and normed on perceived emotional valence, intensity, and social engagement, a new construct. Analyses of ratings collected from 62 young adults (30 females, 32 males; mean age 22 years) revealed the typical quadratic relationship between valence and intensity, i.e., stimuli that are more emotionally charged, whether positively or negatively valenced, are more intense than emotionally-neutral stimuli. Moreover, the results showed significant linear and quadratic relationships between valence and social engagement ratings, indicating that emotionally-charged social scenes were perceived as more engaging than emotionally-neutral social scenes. This new database will facilitate investigations of how people perceive and interpret social and emotional information in everyday interactions, and is offered as a resource to experimenters involved in social and/or emotional processing research.
This article introduces GECO, the Ghent Eye-Tracking Corpus, a monolingual and bilingual corpus of the eyetracking data of participants reading a complete novel. English monolinguals and Dutch–English bilinguals read an entire novel, which was presented in paragraphs on the screen. The bilinguals read half of the novel in their first language, and the other half in their second language. In this article, we describe the distributions and descriptive statistics of the most important reading time measures for the two groups of participants. This large eyetracking corpus is perfectly suited for both exploratory purposes and more directed hypothesis testing, and it can guide the formulation of ideas and theories about naturalistic reading processes in a meaningful context. Most importantly, this corpus has the potential to evaluate the generalizability of monolingual and bilingual language theories and models to the reading of long texts and narratives. The corpus is freely available at http://expsy.ugent.be/downloads/geco.
A central issue in visual and spoken word recognition is the lexical representation of complex words-in particular, whether the lexical representation of complex words depends on semantic transparency: Is a complex verb like understand lexically represented as a whole word or via its base stand, given that its meaning is not transparent from the meanings of its parts? To study this issue, a number of stimulus characteristics are of interest that are not yet available in public databases of German. This article provides semantic association ratings, lexical paraphrases, and vector-based similarity measures for German verbs, measuring (a) the semantic transparency between 1,259 complex verbs and their bases, (b) the semantic relatedness between 1,109 verb pairs with 432 different bases, and (c) the vector-based similarity measures of 846 verb pairs. Additionally, we include the verb regularity of all verbs and two counts of verb family size for 184 base verbs, as well as estimates of age of acquisition and age of reading for 200 verbs. Together with lemma and type frequencies from public lexical databases, all measures can be downloaded along with this article. Statistical analyses indicate that verb family size, morphological complexity, frequency, and verb regularity affect the semantic transparency and relatedness ratings as well as the age of acquisition estimates, indicating that these are relevant variables in psycholinguistic experiments. Although lexical paraphrases, vector-based similarity measures, and semantic association ratings may deliver complementary information, the interrater reliability of the semantic association ratings for each verb pair provides valuable information when selecting stimuli for psycholinguistic experiments.
In this article, we present Procura-PALavras (P-PAL), a Web-based interface for a new European Portuguese (EP) lexical database. Based on a contemporary printed corpus of over 227 million words, P-PAL provides a broad range of word attributes and statistics, including several measures of word frequency (e.g., raw counts, per-million word frequency, logarithmic Zipf scale), morpho-syntactic information (e.g., parts of speech [PoSs], grammatical gender and number, dominant PoS, and frequency and relative frequency of the dominant PoS), as well as several lexical and sublexical orthographic (e.g., number of letters; consonant-vowel orthographic structure; density and frequency of orthographic neighbors; orthographic Levenshtein distance; orthographic uniqueness point; orthographic syllabification; and trigram, bigram, and letter type and token frequencies), and phonological measures (e.g., pronunciation, number of phonemes, stress, density and frequency of phonological neighbors, transposed and phonographic neighbors, syllabification, and biphone and phone type and token frequencies) for {\~{}}53,000 lemmatized and {\~{}}208,000 nonlemmatized EP word forms. To obtain these metrics, researchers can choose between two word queries in the application: (i) analyze words previously selected for specific attributes and/or lexical and sublexical characteristics, or (ii) generate word lists that meet word requirements defined by the user in the menu of analyses. For the measures it provides and the flexibility it allows, P-PAL will be a key resource to support research in all cognitive areas that use EP verbal stimuli. P-PAL is freely available at http://p-pal.di.uminho.pt/tools .
Mean ratings of graphic distinctiveness were obtained for pairs of consonants. The comparisons were between uppercase forms of different consonants, lowercase forms of different consonants, and uppercase vs lowercase forms of the same consonants. The ratings were demonstrated to have satisfactory reliability and to covary moderately well with feature-component measures of letter-pair distinctiveness.
This study aimed to extend the International Affective Picture System (IAPS; Lang, Bradley, {\&} Cuthbert, 2005) norms by obtaining reaction time (RT) normative data for 308 selected photographs. Pictures were presented one at a time for 33, 100, or 250 msec, or under free-time display, to 96 women and 48 men. The participants' task involved assessing the emotional valence of each picture and responding as quickly as possible as to whether it was unpleasant, neutral, or pleasant. RTs provided an index of processing efficiency. The manipulation of display time served to estimate the time course in the valence identification of each picture. Some categories of depicted scenes (e.g., erotica and mutilations) were classified more consistently and efficiently than were others as pleasant or unpleasant. There were minimal differences between men and women. Overall, the present data provide researchers investigating cognition/emotion relationships with an objective criterion to select pictorial stimuli on the basis of RTs. Data for all pictures may be downloaded from brm.psychonomic-journals.org/content/supplemental.
To establish a valid database of vocal emotional stimuli in Mandarin Chinese, a set of Chinese pseudosentences (i.e., semantically meaningless sentences that resembled real Chinese) were produced by four native Mandarin speakers to express seven emotional meanings: anger, disgust, fear, sadness, happiness, pleasant surprise, and neutrality. These expressions were identified by a group of native Mandarin listeners in a seven-alternative forced choice task, and items reaching a recognition rate of at least three times chance performance in the seven-choice task were selected as a valid database and then subjected to acoustic analysis. The results demonstrated expected variations in both perceptual and acoustic patterns of the seven vocal emotions in Mandarin. For instance, fear, anger, sadness, and neutrality were associated with relatively high recognition, whereas happiness, disgust, and pleasant surprise were recognized less accurately. Acoustically, anger and pleasant surprise exhibited relatively high mean f0 values and large variation in f0 and amplitude; in contrast, sadness, disgust, fear, and neutrality exhibited relatively low mean f0 values and small amplitude variations, and happiness exhibited a moderate mean f0 value and f0 variation. Emotional expressions varied systematically in speech rate and harmonics-to-noise ratio values as well. This validated database is available to the research community and will contribute to future studies of emotional prosody for a number of purposes. To access the database, please contact pan.liu@mail.mcgill.ca.
This article introduces a new corpus of eye movements in silent reading—the Russian Sentence Corpus (RSC). Russian uses the Cyrillic script, which has not yet been investigated in cross-linguistic eye movement research. As in every language studied so far, we confirmed the expected effects of low-level parameters, such as word length, frequency, and predictability, on the eye movements of skilled Russian readers. These findings allow us to add Slavic languages using Cyrillic script (exemplified by Russian) to the growing number of languages with different orthographies, ranging from the Roman-based European languages to logographic Asian ones, whose basic eye movement benchmarks conform to the universal comparative science of reading (Share, 2008). We additionally report basic descriptive corpus statistics and three exploratory investigations of the effects of Russian morphology on the basic eye movement measures, which illustrate the kinds of questions that researchers can answer using the RSC. The annotated corpus is freely available from its project page at the Open Science Framework: https://osf.io/x5q2r/.
Words are considered semantically ambiguous if they have more than one meaning and can be used in multiple contexts. A number of recent studies have provided objective ambiguity measures by using a corpus-based approach and have demonstrated ambiguity advantages in both naming and lexical decision tasks. Although the predictive power of objective ambiguity measures has been examined in several alphabetic language systems, the effects in logographic languages remain unclear. Moreover, most ambiguity measures do not explicitly address how the various contexts associated with a given word relate to each other. To explore these issues, we computed the contextual diversity (Adelman, Brown, {\&} Quesada, Psychological Science, 17; 814-823, 2006) and semantic ambiguity (Hoffman, Lambon Ralph, {\&} Rogers, Behavior Research Methods, 45; 718-730, 2013) of traditional Chinese single-character words based on the Academia Sinica Balanced Corpus, where contextual diversity was used to evaluate the present semantic space. We then derived a novel ambiguity measure, namely semantic variability, by computing the distance properties of the distinct clusters grouped by the contexts that contained a given word. We demonstrated that semantic variability was superior to semantic diversity in accounting for the variance in naming response times, suggesting that considering the substructure of the various contexts associated with a given word can provide a relatively fine scale of ambiguity information for a word. All of the context and ambiguity measures for 2,418 Chinese single-character words are provided as supplementary materials.
Sensory experience rating (SER) is a recently developed subjective lexical index that reflects the extent to which a word evokes a sensory and/or perceptual experience in a reader (Juhasz {\&} Yap, 2013; Juhasz, Yap, Dicke, Taylor, {\&} Gullick, 2011). In the present study, SERs for a set of 5,500 Spanish words were collected, which makes this the largest set of norms for SER in the Spanish language to date. Additionally, with the aim of further exploring the implications of this new indicator and its relations with other psycholinguistic variables, a variety of correlational and regression analyses are provided. The results showed that SERs significantly correlated with imageability, age of acquisition, and a number of variables related to perception and emotion. In addition, SERs predicted a significant amount of variance in lexical decision times when other variables were controlled.
Abstract Voice synthesis is a useful method for investigating the communicative role of different acoustic features. Although many text-to-speech systems are available, researchers of human nonverbal vocalizations and bioacousticians may profit from a dedicated simple tool for synthesizing and manipulating natural-sounding vocalizations. Soundgen (https://CRAN.R-project.org/package=soundgen) is an open-source R package that synthesizes nonverbal vocalizations based on meaningful acoustic parameters, which can be specified from the command line or in an interactive app. This tool was validated by comparing the perceived emotion, valence, arousal, and authenticity of 60 recorded human nonverbal vocalizations (screams, moans, laughs, and so on) and their approximate synthetic reproductions. Each synthetic sound was created by manually specifying only a small number of high-level control parameters, such as syllable length and a few anchors for the intonation contour. Nevertheless, the valence and arousal ratings of synthetic sounds were similar to those of the original recordings, and the authenticity ratings were comparable, maintaining parity with the originals for less complex vocalizations. Manipulating the precise acoustic characteristics of synthetic sounds may shed light on the salient predictors of emotion in the human voice. More generally, soundgen may prove useful for any studies that require precise control over the acoustic features of nonspeech sounds, including research on animal vocalizations and auditory perception.
In this article, we present StimulStat – a lexical database for the Russian language in the form of a web application. The database contains more than 52,000 of the most frequent Russian lemmas and more than 1.7 million word forms derived from them. These lemmas and forms are characterized according to more than 70 properties that were demonstrated to be relevant for psycholinguistic research, including frequency, length, phonological and grammatical properties, orthographic and phonological neighborhood frequency and size, grammatical ambiguity, homonymy and polysemy. Some properties were retrieved from various dictionaries and are presented collectively in a searchable form for the first time, the others were computed specifically for the database. The database can be accessed freely at http://stimul.cognitivestudies.ru. {\textcopyright} 2017, Psychonomic Society, Inc.
As the cognitive neuroscience of metaphor has evolved, so too have the theoretical questions of greatest interest. To keep pace with these developments, in the present study we generated a large set of metaphoric and literal sentence pairs ideally suited to addressing the current methodological and conceptual needs of metaphor researchers. In particular, the need has emerged to distinguish metaphors along three dimensions: the grammatical class of their base terms, the sensorimotor features of their base terms, and the syntactic form in which the base terms appear. To meet this need, we generated nominal metaphors (and matched literal sentences) using entity nouns as the base terms, with the intention that they be used in concert with already published sets of predicate metaphors or nominal metaphors using event nouns. Using the results of three norming experiments, we provide 120 pairs of closely matched metaphoric and literal sentences that are characterized along 14 dimensions: 11 at the sentence level (length, frequency, concreteness, familiarity, naturalness, imageability, figurativeness, interpretability, ease of interpretation, valence, and valence judgment reaction time), and three related to the base term (visual, motion, and auditory imagery). These items extend previously published stimuli, filling an extant gap in metaphor research and allowing for tests of new behavioral and neural hypotheses about metaphor.
Fiction is not always accurate, and this has consequences for readers. In laboratory studies, the reading of short stories led participants to produce story errors as facts on a later test of general knowledge (Marsh, Meade, {\&} Roediger, 2003). The present article describes these story stimuli in detail, so that interested researchers will be able to use the stimuli and change them as needed for particular research projects. This article provides instructions for using the stories and suggestions for modifying them; it is a manual for one way of creating suggestibility. The full set of stories and reading comprehension questions may be downloaded fromwww.psychonomic.org/archive/.
Sublexical phonotactic regularities in language have a major impact on language development, as well as on speech processing and production throughout the entire lifespan. To understand the impact of phonotactic regularities on speech and language functions at the behavioral and neural levels, it is essential to have access to oral language corpora to study these complex phenomena in different languages. Yet, probably because of their complexity, oral language corpora remain less common than written language corpora. This article presents the first corpus and database of spoken Quebec French syllables and phones: SyllabO+. This corpus contains phonetic transcriptions of over 300,000 syllables (over 690,000 phones) extracted from recordings of 184 healthy adult native Quebec French speakers, ranging in age from 20 to 97 years. To ensure the representativeness of the corpus, these recordings were made in both formal and familiar communication contexts. Phonotactic distributional statistics (e.g., syllable and co-occurrence frequencies, percentages, percentile ranks, transition probabilities, and pointwise mutual information) were computed from the corpus. An open-access online application to search the database was developed, and is available at www.speechneurolab.ca/syllabo . In this article, we present a brief overview of the corpus, as well as the syllable and phone databases, and we discuss their practical applications in various fields of research, including cognitive neuroscience, psycholinguistics, neurolinguistics, experimental psychology, phonetics, and phonology. Nonacademic practical applications are also discussed, including uses in speech-language pathology.
In this study, we report the validation results of the EU-Emotion Voice Database, an emotional voice database available for scientific use, containing a total of 2,159 validated emotional voice stimuli. The EU-Emotion voice stimuli consist of audio-recordings of 54 actors, each uttering sentences with the intention of conveying 20 different emotional states (plus neutral). The database is organized in three separate emotional voice stimulus sets in three different languages (British English, Swedish, and Hebrew). These three sets were independently validated by large pools of participants in the UK, Sweden, and Israel. Participants' validation of the stimuli included emotion categorization accuracy and ratings of emotional valence, intensity, and arousal. Here we report the validation results for the emotional voice stimuli from each site and provide validation data to download as a supplement, so as to make these data available to the scientific community. The EU-Emotion Voice Database is part of the EU-Emotion Stimulus Set, which in addition contains stimuli of emotions expressed in the visual modality (by facial expression, body language, and social scene) and is freely available to use for academic research purposes.
textcopyright} 2017, The Author(s). Adults need to be able to process infants' emotional expressions accurately to respond appropriately and care for infants. However, research on processing of the emotional expressions of infant faces is hampered by the lack of validated stimuli. Although many sets of photographs of adult faces are available to researchers, there are no corresponding sets of photographs of infant faces. We therefore developed and validated a database of infant faces, which is available via e-mail request. Parents were recruited via social media and asked to send photographs of their infant (0–12 months of age) showing positive, negative, and neutral facial expressions. A total of 195 infant faces were obtained and validated. To validate the images, student midwives and nurses (n = 53) and members of the general public (n = 18) rated each image with respect to its facial expression, intensity of expression, clarity of expression, genuineness of expression, and valence. On the basis of these ratings, a total of 154 images with rating agreements of at least 75{\%} were included in the final database. These comprise 60 photographs of positive infant faces, 54 photographs of negative infant faces, and 40 photographs of neutral infant faces. The images have high criterion validity and good test–retest reliability. This database is therefore a useful and valid tool for researchers.
The rapid expansion of the Internet and the availability of vast repositories of natural text provide researchers with the immense opportunity to study human reactions, opinions, and behavior on a massive scale. To help researchers take advantage of this new frontier, the present work introduces and validates the Evaluative Lexicon 2.0 (EL 2.0)—a quantitative linguistic tool that specializes in the measurement of the emotionality of individuals' evaluations in text. Specifically, the EL 2.0 utilizes natural language to measure the emotionality, extremity, and valence of evaluative reactions and attitudes. The present article describes how we used a combination of 9 million real-world online reviews and over 1,500 participant judges to construct the EL 2.0 and an additional 5.7 million reviews to validate it. To assess its unique value, the EL 2.0 is compared with two other prominent text analysis tools—LIWC and Warriner et al.'s (Behavior Research Methods, 45, 1191–1207, 2013) wordlist. The EL 2.0 is comparatively distinct in its ability to measure emotionality and explains a significantly greater proportion of the variance in individuals' evaluations. The EL 2.0 can be used with any data that involve speech or writing and provides researchers with the opportunity to capture evaluative reactions both in the laboratory and “in the wild.” The EL 2.0 wordlist and normative emotionality, extremity, and valence ratings are freely available from www.evaluativelexicon.com.
The Massive Auditory Lexical Decision (MALD) database is an end-to-end, freely available auditory and production data set for speech and psycholinguistic research, providing time-aligned stimulus recordings for 26,793 words and 9592 pseudowords, and response data for 227,179 auditory lexical decisions from 231 unique monolingual English listeners. In addition to the experimental data, we provide many precompiled listener- and item-level descriptor variables. This data set makes it easy to explore responses, build and test theories, and compare a wide range of models. We present summary statistics and analyses.
In the typical memory conjunction experiment, participants are presented with two "parent" stimulus items (e.g., blackmail and jailbird) that are later recombined to form a "conjunction lure" (e.g., blackbird). This paradigm is an efficient way to test false memories because participants frequently show false recognition for the recombined features of the previously studied stimuli. Two experiments are reported in which normative data for 96 memory conjunction triplets are presented. The first experiment provides descriptive statistics for how often the conjunction triplets show true and false recognition. Due to the variance in the rates of false recognition for the conjunction lure, the second experiment was conducted to help build an understanding of the factors that affect the rate of false recognition of the conjunction lures. Conceptual overlap of the first parent word and the conjunction item predicted false recognition. Digital files containing norms for 96 memory conjunction triplets may be downloaded from www.psychonomic.org/archive.
An adult language corpus of spoken Hong Kong Cantonese (HKCAC) has recently been developed consisting of spontaneous speech recorded from phone-in programs and forums on the radio in Hong Kong. The database represents the speech of a total of sixty-nine speakers in addition to the program hosts, and has approximately 170, 000 characters. It is believed that HKCAC will be of great value to linguists who are interested in studying Cantonese, and speech therapists and educators who work with the Cantonese speaking population. A search engine with a user-friendly interface has also been developed by using FileMaker Pro 4.0 (Chinese version). Apart from the basic frequency information and the display of search results in KWAL (Key Word And Line) format, the search engine also allows users to search for various phonetic realizations of a particular character or the set of characters associated with a particular syllable. The content and structure of the corpus, and the overall architecture as well as the technical aspects of the search engine are described. Search procedures are illustrated with examples. The paper ends with a discussion of the future development of HKCAC. {\textcopyright} 2001 John Benjamins Publishing Company.
Recent studies have shown that word frequency estimates obtained from films and television subtitles are better to predict performance in word recognition experiments than the traditional word frequency estimates based on books and newspapers. In this study, we present a subtitle-based word frequency list for Spanish, one of the most widely spoken languages. The subtitle frequencies are based on a corpus of 41M words taken from contemporary movies and TV series (screened between 1990 and 2009). In addition, the frequencies have been validated by correlating them with the RTs from two megastudies involving 2,764 words each (lexical decision and word naming tasks). The subtitle frequencies explained 6{\%} more of the variance than the existing written frequencies in lexical decision, and 2{\%} extra in word naming.
This article presents the Provo Corpus, a corpus of eye-tracking data with accompanying predictability norms. The predictability norms for the Provo Corpus differ from those of other corpora. In addition to traditional cloze scores that estimate the predictability of the full orthographic form of each word, the Provo Corpus also includes measures of the predictability of the morpho-syntactic and semantic information for each word. This makes the Provo Corpus ideal for studying predictive processes in reading. Some analyses using these data have previously been reported elsewhere (Luke {\&} Christianson, 2016). The Provo Corpus is available for download on the Open Science Framework, at https://osf.io/sjefs .
We describe the Multilanguage Written Picture Naming Dataset. This gives trial-level data and time and agreement norms for written naming of the 260 pictures of everyday objects that compose the colorized Snodgrass and Vanderwart picture set (Rossion {\&} Pourtois in Perception, 33, 217–236, 2004). Adult participants gave keyboarded responses in their first language under controlled experimental conditions (N = 1,274, with subsamples responding in Bulgarian, Dutch, English, Finnish, French, German, Greek, Icelandic, Italian, Norwegian, Portuguese, Russian, Spanish, and Swedish). We measured the time to initiate a response (RT) and interkeypress intervals, and calculated measures of name and spelling agreement. There was a tendency across all languages for quicker RTs to pictures with higher familiarity, image agreement, and name frequency, and with higher name agreement. Effects of spelling agreement and effects on output rates after writing onset were present in some, but not all, languages. Written naming therefore shows name retrieval effects that are similar to those found in speech, but our findings suggest the need for cross-language comparisons as we seek to understand the orthographic retrieval and/or assembly processes that are specific to written output.
Recent research on anagram solution has produced two original findings. First, it has shown that a new bigram frequency measure called top rank, which is based on a comparison of summed bigram frequencies, is an important predictor of anagram difficulty. Second, it has suggested that the measures from a type count are better than token measures at predicting anagram difficulty. Testing these hypotheses has been difficult because the computation of the bigram statistics is difficult. We present a program that calculates bigram measures for two-to nine-letter words. We then show how the program can be used to compare the contribution of top rank and other bigram frequency measures derived from both a token and a type count. Contrary to previous research, we report that type measures are not better at predicting anagram solution times and that top rank is not the best predictor of anagram difficulty. Lastly we use this program to show that type bigram frequencies are not as good as token bigram frequencies at predicting word identification reaction time.
Rebus puzzles and compound remote associate problems have been successfully used to study problem solving. These problems are physically compact, often can be solved within short time limits, and have unambiguous solutions, and English versions have been normed for solving rates and levels of difficulty. Many studies on problem solving with sudden insight have taken advantage of these features in paradigms that require many quick solutions (e.g., solution priming, visual hemifield presentations, electroencephalography, fMRI, and eyetracking). In order to promote this vein of research in Italy, as well, we created and tested Italian versions of both of these tests. The data collected across three studies yielded a pool of 88 rebus puzzles and 122 compound remote associate problems within a moderate range of difficulty. This article provides both sets of problems with their normative data, for use in future research.
This article accompanies the archiving by the Pychonomic Society of the Toglia and Battig (1978) semantic word norms. Herein are outlined the various phases of the project, as well as the challenges that were faced in staying the course during the labor-intensive development of the norms. An examination of the number of citations of this set of norms over the years demonstrates a stable employment of these norms by investigators in many fields. Indeed, a concluding section details the wide range of research topics that have been studied with the use of this extensive set of word ratings. The complete Toglia and Battig article and norms may be downloaded as supplemental materials for this article from brm.psychonomic-journals.org/content/supplemental.
We present word prevalence data for 61,858 English words. Word prevalence refers to the number of people who know the word. The measure was obtained on the basis of an online crowdsourcing study involving over 220,000 people. Word prevalence data are useful for gauging the difficulty of words and, as such, for matching stimulus materials in experimental conditions or selecting stimulus materials for vocabulary tests. Word prevalence also predicts word processing times, over and above the effects of word frequency, word length, similarity to other words, and age of acquisition, in line with previous findings in the Dutch language.
Numerous studies in psychology, cognitive neuroscience and psycholinguistics have used pictures of objects as stimulus materials. Currently, authors engaged in cross-linguistic work or wishing to run parallel studies at multiple sites where different languages are spoken must rely on rather small sets of black-and-white or colored line drawings. These sets are increasingly experienced as being too limited. Therefore, we constructed a new set of 750 colored pictures of concrete concepts. This set, MultiPic, constitutes a new valuable tool for cognitive scientists investigating language, visual perception, memory and/or attention in monolingual or multilingual populations. Importantly, the MultiPic databank has been normed in six different European languages (British English, Spanish, French, Dutch, Italian and German). All stimuli and norms are freely available at http://www.bcbl.eu/databases/multipic.
Iconicity – the correspondence between form and meaning – may help young children learn to use new words. Early‐learned words are higher in iconicity than later learned words. However, it remains unclear what role iconicity may play in actual language use. Here, we ask whether iconicity relates not just to the age at which words are acquired, but also to how frequently children and adults use the words in their speech. If iconicity serves to bootstrap word learning, then we would expect that children should say highly iconic words more frequently than less iconic words, especially early in development. We would also expect adults to use iconic words more often when speaking to children than to other adults. We examined the relationship between frequency and iconicity for approximately 2000 English words. Replicating previous findings, we found that more iconic words are learned earlier. Moreover, we found that more iconic words tend to be used more by younger children, and adults use more iconic words when speaking to children than to other adults. Together, our results show that young children not only learn words rated high in iconicity earlier than words low in iconicity, but they also produce these words more frequently in conversation – a pattern that is reciprocated by adults when speaking with children. Thus, the earliest conversations of children are relatively higher in iconicity, suggesting that this iconicity scaffolds the production and comprehension of spoken language during early development.
ABSTRACTIn brain and behaviour, gustation, and olfaction are closely linked to emotional processing. This paper shows that similarly, words associated with taste and smell, such as “pungent” and “delicious”, are on average more emotionally valenced than words associated with the other senses, such as “beige” (visual) and “echoing” (auditory). Moreover, taste and smell words occur more frequently in emotionally valenced phrases, for example, “fragrant” modifies more emotionally valenced nouns (“fragrant kiss”) than the visual adjective “yellow” (“yellow house”). It is argued that taste and smell words form an affectively loaded part of the English lexicon. Taste and smell words are also shown to be more emotionally flexible in that words such as “sweet” can be combined with both good and bad nouns (“sweet delight” versus “sweet disaster”), much more so than is the case for sensory words for the other modalities. The paper discusses implications for theories of embodied language understanding.
This paper presents an overview of a project that aims at creating a representative Catalogue of cross-linguistically recurrent semantic shifts1 in the languages of the world and at implementing this Catalogue in the form of a searchable computer database2. Such a catalogue is useful in several theoretical and methodological respects. First of all, both universal and language-specific semantic shifts can be considered a window onto human cognitive mechanisms operative in the domain of linguistic conceptualization. Second, the catalogue provides a rich empirical basis for the study of genetic and areal tendencies in semantic change, as well as in polysemy patterns. Another potential application is historical reconstruction, since the catalogue gives evidence for attested paths of diachronic semantic evolution. In section 1 we outline the general concept of the Catalogue of Semantic Shifts. Section 2 describes the design of the computer database. In section 3 we discuss some problematic points that we have faced while working on the Catalogue. The next three sections demonstrate how the Catalogue can be used in linguistic research and present an analysis of three selected issues, namely semantic shifts in the domain of dimension (section 4), motivation strategies in the domain of folk biology (section 5), and euphemization as a mechanism of semantic change (section 6).
This article describes the preparation, recording and orthographic transcription of a new speech corpus, the Nijmegen Corpus of Casual French (NCCFr). The corpus contains a total of over 36 h of recordings of 46 French speakers engaged in conversations with friends. Casual speech was elicited during three different parts, which together provided around 90 min of speech from every pair of speakers. While Parts 1 and 2 did not require participants to perform any specific task, in Part 3 participants negotiated a common answer to general questions about society. Comparisons with the ESTER corpus of journalistic speech show that the two corpora contain speech of considerably different registers. A number of indicators of casualness, including swear words, casual words, verlan, disfluencies and word repetitions, are more frequent in the NCCFr than in the ESTER corpus, while the use of double negation, an indicator of formal speech, is less frequent. In general, these estimates of casualness are constant through the three parts of the recording sessions and across speakers. Based on these facts, we conclude that our corpus is a rich resource of highly casual speech, and that it can be effectively exploited by researchers in language science and technology. {\textcopyright} 2009 Elsevier B.V. All rights reserved.
We constructed a corpus of digitized texts containing about 4{\%} of all books ever printed. Analysis of this corpus enables us to investigate cultural trends quantitatively. We survey the vast terrain of 'culturomics,' focusing on linguistic and cultural phenomena that were reflected in the English language between 1800 and 2000. We show how this approach can provide insights about fields as diverse as lexicography, the evolution of grammar, collective memory, the adoption of technology, the pursuit of fame, censorship, and historical epidemiology. Culturomics extends the boundaries of rigorous quantitative inquiry to a wide array of new phenomena spanning the social sciences and the humanities.
The Corpus of Contemporary American English is the first large, genre-balanced corpus of any language, which has been designed and constructed from the ground up as a ‘monitor corpus', and which can be used to accurately track and study recent changes in the language. The 400 million words corpus is evenly divided between spoken, fiction, popular magazines, newspapers, and academic journals. Most importantly, the genre balance stays almost exactly the same from year to year, which allows it to accurately model changes in the ‘real world'. After discussing the corpus design, we provide a number of concrete examples of how the corpus can be used to look at recent changes in English, including morphology (new suffixes –friendly and –gate), syntax (including prescriptive rules, quotative like, so not ADJ, the get passive, resultatives, and verb complementation), semantics (such as changes in meaning with web, green, or gay), and lexis––including word and phrase frequency by year, and using the corpus architecture to produce lists of all words that have had large shifts in frequency between specific historical periods.
Word processing studies increasingly make use of regression analyses based on large numbers of stimuli (the so-called megastudy approach) rather than experimental designs based on small factorial designs. This requires the availability of word features for many words. Following similar studies in English, we present and validate ratings of age of acquisition and concreteness for 30,000 Dutch words. These include nearly all lemmas language researchers are likely to be interested in. The ratings are freely available for research purposes.
This paper reports on the TIGER Treebank, a corpus of currently 40,000 syntactically annotated German newspaper sentences. We describe what kind of information is encoded in the treebank and introduce the different representation formats that are used for the annotation and exploitation of the treebank. We explain the different methods used for the annotation: interactive annotation, using the tool ANNOTATE, and LFG parsing. Furthermore, we give an account of the annotation scheme used for the TIGER treebank. This scheme is an extended and improved version of the NEGRA annotation scheme and we illustrate in detail the linguistic extensions that were made concerning the annotation in the TIGER project. The main differences are concerned with coordination, verb-subcategorization, expletives as well as proper nouns. In addition, the paper also presents the query tool TIGERSearch that was developed in the project to exploit the treebank in an adequate way. We describe the query language which was designed to facilitate a simple formulation of complex queries; furthermore, we shortly introduce TIGER in, a graphical user interface for query input. The paper concludes with a summary and some directions for future work.
The York–Toronto–Helsinki Parsed Corpus of Old English Prose (YCOE) is a 1.5 million-word syntactically annotated corpus of Old English prose texts. It was produced at the University of York, UK, between 2000 and 2003, by Ann Taylor, Anthony Warner, Susan Pintzuk and Frank Beths, with a grant from the English Arts and Humanities Research Board (B/RG/AN5907/APN9528). The YCOE is part of the English Parsed Corpora Series. It was the third historical corpus to be completed in this format, and uses the same kind of annotation scheme as its sister corpora, the Penn–Helsinki Parsed Corpus of Middle English II (PPCME2), the York–Helsinki Parsed Corpus of Old English Poetry and the Penn–Helsinki Parsed Corpus of Early Modern English. Two other corpora in the series, the parsed version of the Corpus of Early English Correspondence and the Penn Parsed Corpus of Modern British English, are currently under construction (at the University of York, UK, in cooperation with the University of Helsinki, Finland, and the University of Pennsylvania, USA, respectively).
Developmental differences in name agreement, familiarity, and visual complexity in response to line drawings of common objects were obtained from children and adults. Sixty-one pictures were taken from the Peabody Picture Vocabulary Test-Revised and 259 pictures were taken from the set normed for adults by Snodgrass and Vanderwart (1980). Although there were some differences between the two sets of pictures, the present results replicated the relative independence of these three measures, which was reported by Snodgrass and Vanderwart for adults. Children and adults showed substantial agreement on the names of the pictures. Although the children's ratings were lower on all measures, the differences were trivial for most pictures. We concluded that judgments of familiarity, complexity, and the names of line drawings of common objects are based primarily on information processing accomplished prior to age 7.
According to recent embodied cognition theories, mental concepts are represented by modality-specific sensory-motor systems. Much of the evidence for modality-specificity in conceptual processing comes from the property-verification task. When applying this and other tasks, it is important to select items based on their modality-exclusivity. We collected modality ratings for a set of 387 properties, each of which was paired with two different concepts, yielding a total of 774 concept-property items. For each item, participants rated the degree to which the property could be experienced through five perceptual modalities (vision, audition, touch, smell, and taste). Based on these ratings, we computed a measure of modality exclusivity, the degree to which a property is perceived exclusively through one sensory modality. In this paper, we briefly sketch the theoretical background of conceptual knowledge, discuss the use of the property-verification task in cognitive research, provide our norms and statistics, and validate the norms in a memory experiment. We conclude that our norms are important for researchers studying modality-specific effects in conceptual processing.
[Correction Notice: An Erratum for this article was reported in Vol 49(3) of Behavior Research Methods (see record 2017-21867-030). In the original article, there was an error in the second sentence of Footnote 3 on page 175. The corrected sentence is present in the erratum.] The body-shape-related stimuli used in most body-image studies have several limitations (e.g., a lack of pilot validation procedures and the use of non-body-shape-related control/neutral stimuli). We therefore developed a database of 61 computer-generated body-only pictures of women, wherein bodies were methodically manipulated in terms of fatness versus thinness. Eighty-two young women assessed the pictures' attractiveness, beauty, harmony (valence ratings), and body shape (assessed on a thinness/fatness axis), providing normative data for valence and body shape ratings. First, stimuli manipulated for fatness versus thinness conveyed comparable emotional intensities regarding the valence and body shape ratings. Second, different subcategories of stimuli were obtained on the basis of variations in body shape and valence judgments. Fat and thin bodies were distributed into several subcategories depending on their valence ratings, and a subcategory containing stimuli that were neutral in terms of valence and body shape was identified. Interestingly, at a descriptive level, the thinness/fatness manipulations of the bodies were in a curvilinear relationship with the valence ratings: Thin bodies were not only judged as positive, but also as negative when their estimated body mass indexes (BMIs) decreased too much. Finally, convergent validity was assessed by exploring the impacts of body-image-related variables (BMI, thin-ideal internalization, and body dissatisfaction) on participants' judgments of the bodies. Valence judgments, but not body shape judgments, were influenced by the participants' levels of thin-ideal internalization and body dissatisfaction. Participants' BMIs did not significantly influence their judgments. Given these findings, this database contains relevant material that can be used in various fields, primarily for studies of body-image disturbance or eating disorders. (PsycINFO Database Record (c) 2017 APA, all rights reserved) (Source: journal abstract)
Reading involves a process of matching an orthographic input with stored representations in lexical memory. The masked priming paradigm has become a standard tool for investigating this process. Use of existing results from this paradigm can be limited by the precision of the data and the need for cross-experiment comparisons that lack normal experimental controls. Here, we present a single, large, high-precision, multicondition experiment to address these problems. Over 1,000 participants from 14 sites responded to 840 trials involving 28 different types of orthographically related primes (e.g., castfe-CASTLE) in a lexical decision task, as well as completing measures of spelling and vocabulary. The data were indeed highly sensitive to differences between conditions: After correction for multiple comparisons, prime type condition differences of 2.90 ms and above reached significance at the 5{\%} level. This article presents the method of data collection and preliminary findings from these data, which included replications of the most widely agreed-upon differences between prime types, further evidence for systematic individual differences in susceptibility to priming, and new evidence regarding lexical properties associated with a target word's susceptibility to priming. These analyses will form a basis for the use of these data in quantitative model fitting and evaluation and for future exploration of these data that will inform and motivate new experiments.
The extent to which processing words involves breaking them down into smaller units or morphemes or is the result of an interactive activation of other units, such as meanings, letters, and sounds (e.g., dis-agree-ment vs. disagreement), is currently under debate. Disentangling morphology from phonology and semantics is often a methodological challenge, because orthogonal manipulations are difficult to achieve (e.g., semantically unrelated words are often phonologically related: casual-casualty and, vice versa, sign-signal). The present norms provide a morphological classification of 3,263 suffixed derived words from two widely spoken languages: English (2,204 words) and Spanish (1,059 words). Morphologically complex words were sorted into four categories according to the nature of their relationship with the base word: phonologically transparent (friend-friendly), phonologically opaque (child-children), semantically transparent (habit-habitual), and semantically opaque (event-eventual). In addition, ratings were gathered for age of acquisition, imageability, and semantic distance (i.e., the extent to which the meaning of the complex derived form could be drawn from the meaning of its base constituents). The norms were completed by adding values for word frequency; word length in number of phonemes, letters, and syllables; lexical similarity, as measured by the number of neighbors; and morphological family size. A series of comparative analyses from the collated ratings for the base and derived words were also carried out. The results are discussed in relation to recent findings.