1396 norm sets
The Online Malay Language Corpus-based Lexical Database for Primary Schools discussed in this paper is a long-term research project by the School of Educational Studies, Universiti Sains Malaysia. Based on a corpus of Malay language textbooks used in primary schools, the project aims to develop a lexical database of Malay words commonly encountered by elementary school children in Malaysia. Available online, the database has an interactive interface that allows users to search in real-time primary linguistic features such as word frequency, word length, phoneme length, number and type of syllables, as well as word category. The database has proven to be a useful resource for both researchers and practitioners who use it to identify linguistically and culturally appropriate sets of word stimuli for material development, language assessment, language teaching, and language remediation. These applications facilitate and promote evidence-based teaching and research practices pertaining to reading acquisition in the Malay language. The system has undergone initial validation and is continuously being revised and expanded to improve on its usability, readability, and accuracy.
These materials consist of lists of English collocations along with various associated measures. All these collocations are composed of two words. The most common structures are N-N, Adj-N, and V-‘Any’. There are two sets of core collocations, one for valence (N = 121) and one for arousal (N = 124). Some lists include calibrator and control items. The two core sets overlap substantially. The ratings of valence and arousal for whole collocations were crowd-sourced using Amazon Mechanical Turk (AMT) (https://www.mturk.com/), The ratings of the constituent words stem from a list compiled by Warriner, Kuperman, and Brysbaert (2013). The list itself is available at: http://crr.ugent.be/archives/1003 For further details see the main text in the Elsevier journal of applied linguistics, 'System'. The article's title is: 'Ratings of the emotional valence and arousal of collocations and their constiituent words: How can they be useful in L2 vocabulary research?'. Associated with each core set of collocations is a smaller set of collocations having ratings from both AMT and Warriner et al. These matched, ‘overlapping’ ratings were used to assess the reliability of the new AMT ratings. In the main lists of collocations these ‘overlappers’ are given in red italics. Additionally, there are lists of regression residuals. The closer a residual is to zero, the more accurately rated collocations valence (or arousal) was predicted by the valence (or arousal) ratings of the mean of the constituent word ratings. <b> </b>Important abbreviations used in the spreadsheet are: AMT = Amazon Mechanical Turk; CW = Constituent word; Geo.mean = Geometrical mean; Harm.mean = Harmonic mean; Most valenced = The CW rating that is the furthest from 5 (i.e., neutral) either toward 1 or toward 9; SD = The standard deviation of the individual AMT ratings obtained for a given collocation; WKB = Warriner et al. (2013); NA = not available. NA was used in place of values (e.g., CW ratings) that could not be found in the list of WKB. This abbreviation was chosen because the R functions used in the studies can handle datasets that include NA in place of a missing value. For instance, the appropriate calls in base R for calculating Spearman’s and Pearson’s correlations<i> </i>between the variables x and y, when NAs are present, are: <i>cor(x, y, method = "s", use = "pairwise.complete.obs")</i> and<i> cor(x, y, method = "p", use = "pairwise.complete.obs"). </i>The R functions of Wilcox (2012) that were used handle missing values even more automatically. For example, when Wilcox’s R functions are installed (https://dornsife.usc.edu/labs/rwilcox/software/), the call <i>corb(x,y, corfun = spear, nboot = 20000)</i> gives a bootstrap 95% confidence interval for Spearman’s correlation. The call for a median-based linear regression method would be: <i>tsreg(x, y)</i>, where x and y are the independent and the dependent variables, respectively. <b>References</b> Warriner, A., Kuperman, V., & Brysbaert, M. (2013). Norms of valence, arousal, and dominance for 13,915 English lemmas. <i>Behavior Research Methods, 45</i>, 1191–11207. https://doi.org/10.3758/s13428-012-0314-x Wilcox, R. (2012). <i>Modern statistics for the social and behavioural sciences</i>. Boca Raton, FL: CRC Press. <i> </i>
Over the last 40 years, object recognition studies have moved from using simple line drawings, to more detailed illustrations, to more ecologically valid photographic representations. Researchers now have access to various stimuli sets, however, existing sets lack the ability to independently manipulate item format, as the concepts depicted are unique to the set they derive from. To enable such comparisons, Rossion and Pourtois (2004) revisited Snodgrass and Vanderwart's (1980) line drawings and digitally re-drew the objects, adding texture and shading. In the current study, we took this further and created a set of stimuli that showcase the same objects in photographic form. We selected six photographs of each object (three color/three grayscale) and collected normative data and RTs. Naming accuracy and agreement was high for all photographs and appeared to steadily increase with format distinctiveness. In contrast to previous data patterns for drawings, naming agreement (H values) did not differ between grey and color photographs, nor did familiarity ratings. However, grey photographs received significantly lower mental imagery agreement and visual complexity scores than color photographs. This suggests that, in comparison to drawings, the ecological nature of photographs may facilitate deeper critical evaluation of whether they offer a good match to a mental representation. Color may therefore play a more vital role in photographs than in drawings, aiding participants in judging the match with their mental representation. This new photographic stimulus set and corresponding normative data provide valuable materials for a wide range of experimental studies of object recognition.
This study aimed at providing subjective frequency and imageability norms for a sample of 1,760 monosyllabic French words and thereby, increasing the pool of normative data available for research in cognitive science and language processing. The results indicate that the reliability of the estimates is high, with coefficients ranging between.93 and.99 for the frequency and imageability ratings. External validity was investigated by calculating correlations with ratings drawn from all similar studies and for which the number of shared items was sufficient. These coefficients vary between.73 and.88 for subjective frequency and between.64 and.97 for imageability. The correlation between subjective frequency and imageability in the present study was significant and relatively high (r =.64). The implications of these results for the selection of experimental stimuli for research are discussed.
This paper describes the creation of a normed lexicon of 150 pairs of English and Italian idioms annotated by translatability level (Beck, 2020). The dataset was created through the implementation of a cross-linguistic norming study conducted online via a novel combination of tools and design, which was validated by data reliability measurement. The lexicon contains information with respect to two groups of idiom variables (Hubers, Cucchiarini, Strik, & Dijkstra, 2019): Experience-Based Variables (“EBVs”: familiarity, meaningfulness and objective knowledge) and Content-Based Variables (“CBVs”: literal plausibility, decomposability and transparency). These variables are particularly relevant for the analysis of ambiguous contexts where there is an interaction between the literal and figurative meaning of an idiom (Wagner, 2021). In particular, the sum of the mean ratings obtained for the three CBVs provides the Potential Idiomatic Ambiguity (“PIA”) index of each idiom: the higher the PIA, the more likely the idiom should be to occur in ambiguous contexts. This index can therefore be exploited in the setup of future psycholinguistic experiments to verify the relationship between the internal distribution of idiom features and ambiguous contexts. Moreover, the lexicon is a valuable research tool per se, since it contributes to bridging the gap in cross-linguistic research in idiom norming studies (Nordmann, Cleland, & Bull, 2014; Pastor, 2021), enabling systematic comparative analyses and enhancing understanding of the relationships among the examined variables.
Psycholinguistic research shows that numerous variables influence language comprehension and production.<b> </b>Despite the importance of norming words for neutralizing lexical confounding factors, facilitating cross-linguistic comparative research, and ensuring the replicability of psychological experiments, such lexical resources remain disproportionately scarce for minority and under-resourced languages. This dataset presents NORGAL, a comprehensive database comprising normative ratings for emotional valence, arousal, concreteness, imageability and subjective age of acquisition (AoA) of 1,585 Galician words. Highly frequent words – nouns and adjectives in particular – were extracted from Reference Corpus of Current Galician (CORGA) and normed on a 9-point Likert scale (except for AoA) by bilingual Galician-Spanish students from two University of Vigo campuses. This study also presents comparative analyses against established Spanish normative databases, ensuring strict methodological control and confirming the consistency of our stimulus ratings across studies. Finally, primary limitations associated with normative research in acquiring psycholinguistic indices in this specific context are explained, while paving the way for future avenues of application and expansion. In sum, NORGAL can serve as a valuable foundation for advancing experimental and psycholinguistic research in this language.
Emerging opportunities for global big data sharing and analysis make it vital to adapt and develop neuropsychological tests in a way that takes account of psycholinguistic variables that may vary between cultures and thus differentially bias participants' responses. Robust regional norms will be essential to the identification of these variables. We have built three normative corpuses with Argentinean scores for concept, image and features variables. In the current presentation we will first describe each of them, then illustrate the way they can be used in test construction and adaptation, and third establish some comparisons between languages. three Argentinean normative databases: 1) Expanded norms for 400 experimental pictures (Manoiloff et al., 2010); 2) Semantic and latency times norms (Martínez-Cuitiño et al., 2015); 3) Semantic feature production norms (Vivas et al., 2017). Data analysis: a) Pearson correlations were calculated to compare Spanish, French and English psycholinguistic norms for name agreement (NA), image agreement, familiarity, visual complexity and image variability; b) A geometric vector comparison technique was performed between English and Spanish feature lists for 200 concepts. a) NA proved the most dependent on language and/or culture, since it obtained the most diverse values (Min = 18, Max = 100) and the lowest correlations between the Argentinean study and previous studies in Spanish (r =.54), French (r =.37) and English (r =.30); b) Similarities between the core components of semantic representations in Spanish and English were observed in most of the concepts, although some relevant features differed between languages in certain concepts (eg. football was the most relevant feature for the concept BALL in Argentine but did not appear in English norms). Feature concordance was higher for living than for non-living concepts (z = -4.638; p <.001). Our results show some coincidences between languages but also some important differences in relevant psycholinguistic variables which make it essential to consider local normative data in order to select the optimal stimuli to build neuropsychological tests with equivalent properties across languages.
We provide psycholinguistic norms for a new set of 160 French idiomatic expressions and 160 proverbs: knowledge, predictability, literality, compositionality, subjective and objective frequency, familiarity, age of acquisition (AoA) and length. Different analyses (reliability, descriptive statistics and correlations) performed on the norms are reported and discussed. The norms can be downloaded as Supplemental Material.
This study investigates the potential of large language models (LLMs) to estimate the familiarity of words and multi-word expressions (MWEs). We validated LLM estimates for isolated words using existing human familiarity ratings and found strong correlations. LLM familiarity estimates performed even better in predicting lexical decision and naming performance in megastudies than the best available word frequency measures. We then applied LLM estimates to MWEs, also finding their effectiveness in measuring familiarity for these expressions. We have created a list of more than 400,000 English words and MWEs with LLM-generated familiarity estimates, which we hope will be a valuable resource for researchers. There is also a cleaned-up list of nearly 150,000 entries, excluding lesser-known stimuli, to streamline stimulus selection. Our findings highlight the advantages of LLM-based familiarity estimates, including their better performance than traditional word frequency measures (particularly for predicting word recognition accuracy), their ability to generalize to MWEs, availability for large lists of words, and ease of obtaining new estimates for all types of stimuli.
In this paper, we introduce the resources that we developed for Turkish\ndependency parsing, which include a novel manually annotated treebank (BOUN\nTreebank), along with the guidelines we adopted, and a new annotation tool\n(BoAT). The manual annotation process we employed was shaped and implemented by\na team of four linguists and five Natural Language Processing (NLP)\nspecialists. Decisions regarding the annotation of the BOUN Treebank were made\nin line with the Universal Dependencies (UD) framework as well as our recent\nefforts for unifying the Turkish UD treebanks through manual re-annotation. To\nthe best of our knowledge, BOUN Treebank is the largest Turkish treebank. It\ncontains a total of 9,761 sentences from various topics including biographical\ntexts, national newspapers, instructional texts, popular culture articles, and\nessays. In addition, we report the parsing results of a state-of-the-art\ndependency parser obtained over the BOUN Treebank as well as two other\ntreebanks in Turkish. Our results demonstrate that the unification of the\nTurkish annotation scheme and the introduction of a more comprehensive treebank\nlead to improved performance with regard to dependency parsing.\n
We present a new set of subjective age-of-acquisition (AoA) ratings for 299 words (158 nouns, 141 verbs) in 25 languages from five language families (Afro-Asiatic: Semitic languages; Altaic: one Turkic language: Indo-European: Baltic, Celtic, Germanic, Hellenic, Slavic, and Romance languages; Niger-Congo: one Bantu language; Uralic: Finnic and Ugric languages). Adult native speakers reported the age at which they had learned each word. We present a comparison of the AoA ratings across all languages by contrasting them in pairs. This comparison shows a consistency in the orders of ratings across the 25 languages. The data were then analyzed (1) to ascertain how the demographic characteristics of the participants influenced AoA estimations and (2) to assess differences caused by the exact form of the target question (when did you learn vs. when do children learn this word); (3) to compare the ratings obtained in our study to those of previous studies; and (4) to assess the validity of our study by comparison with quasi-objective AoA norms derived from the MacArthur-Bates Communicative Development Inventories (MB-CDI). All 299 words were judged as being acquired early (mostly before the age of 6 years). AoA ratings were associated with the raters' social or language status, but not with the raters' age or education. Parents reported words as being learned earlier, and bilinguals reported learning them later. Estimations of the age at which children learn the words revealed significantly lower ratings of AoA. Finally, comparisons with previous AoA and MB-CDI norms support the validity of the present estimations. Our AoA ratings are available for research or other purposes.
ABSTRACT The interplay between emotion and language has drawn increasing attention from researchers in various fields, such as psycholinguistics, neurolinguistics, and artificial intelligence. The classification of emotional words is crucial to experimental studies, as it serves as a fundamental step for unveiling the complex relationship between emotion and language. Hence, the current study introduces affective and psycholinguistic norms for 1200 two‐character Chinese words. The affective variables, including emotional prototypicality (EmoPro), valence, and arousal, as well as psycholinguistic variables (abstractness and familiarity), were rated by native speakers of Chinese using a 7‐point Likert scale. This set of norms, as far as we know, is the first one that provides words’ EmoPro together with other affective and psycholinguistic variables rated by the same group of participants. The results of inter‐rater reliability and correlation tests showed that this set of norms had good reliability and validity. The correlation between EmoPro and arousal was significant and strong. Valence and arousal exhibited an asymmetric U‐shaped relationship, with negative words being rated higher in arousal than positive ones. We also identified a quadratic relationship between EmoPro and valence, indicating that more prototypical emotion‐label words tend to trigger more extreme valence ratings. EmoPro significantly and positively correlated with abstractness, indicating that prototypical emotion‐label words tend to be more abstract. In addition, familiarity and word frequency predict lexical decision performance, whereas EmoPro does not after controlling for other variables. The present set of Chinese norms supports material selection in experimental research, enables cross‐linguistic comparisons, and has potential applications in natural language processing.
In this paper, the authors provide a data report that describes an original dataset named French Affective Images of Climate Change (FAICC) database. The main objective is to provide tools for CC assessment. Images are rated by a sample of non-experts according to three variables: relevance to CC, arousal, and emotional valence. The database provides for each image an identification number, the mean rating and standard deviation of ratings for relevance, arousal and valence, respectively.
The Affective Norms for English Words (ANEW; Bradley & Lang, 1999) scale is a widely used instrument for valence and arousal response in English. A person whose first language is American Sign Language (ASL) might process the English emotion words differently. We hypothesized that ASL users might provide different valence and arousal ratings for emotion words in ASL, and a separate normative database might be necessary for this population. Forty-two Deaf adult signers completed ratings for the English and ASL conditions. Results showed that the rating for the arousal were similar for both conditions. However, the valence ratings were different, which could be explained by the different word frequency among the ASL users. This raises a need to create a separate valence rating normative database in ASL.
We present the Chinese Lexical Database (CLD). The CLD is a new large-scale lexical database for Mandarin Chinese that provides over 150 descriptive and lexical-distributional variables for more than 30,000 words in simplified Chinese. The information in the CLD can be used for the construction of experimental stimuli and the analysis of experimental data in psycholinguistic research on simplified Chinese. The CLD can be downloaded for free, and an online search interface is provided at http://www.chineselexicaldatabase.com.
We present the Chinese Lexical Database 2.1 (CLD 2.1). The CLD is a new large-scale lexical database for Mandarin Chinese that provides over 260 descriptive and lexical-distributional variables for more than 48,644 words in simplified Chinese. The information in the CLD can be used for the construction of experimental stimuli and the analysis of experimental data in psycholinguistic research on simplified Chinese. The CLD can be downloaded for free, and an online search interface is provided at http://www.chineselexicaldatabase.com.
In this article we present the RST Spanish Treebank, the first corpus annotated with rhetorical relations for this language. We describe the characteristics of the corpus, the annotation criteria, the annotation procedure, the inter-annotator agreement, and other related aspects. Moreover, we show the interface that we have developed to carry out searches over the corpus' annotated texts.
Emotional databases are important tools to study emotion recognition and their effects on various cognitive processes. Since, well-standardized large-scale emotional expression database is not available in India, we evaluated Radboud faces database (RaFD)-a freely available database of emotional facial expressions of adult Caucasian models, for Indian sample. Using the pictures from RaFD, we investigated the similarity and differences in self-reported ratings on emotion recognition accuracy as well as parameters of valence, clarity, genuineness, intensity and arousal of emotional expression, by following the same rating procedure as used for the validation of RaFD. We also systematically evaluated the universality hypothesis of emotion perception by analyzing differences in accuracy and ratings for different emotional parameters across Indian and Dutch participants. As the original Radboud database lacked arousal rating, we added this as a emotional parameter along with all other parameters. The results show that the overall accuracy of emotional expression recognition by Indian participants was high and very similar to the ratings from Dutch participants. However, there were significant cross-cultural differences in classification of emotion categories and their corresponding parameters. Indians rated certain expressions comparatively more genuine, higher in valence, and less intense in comparison to original Radboud ratings. The misclassifications/ confusion for specific emotional categories differed across the two cultures indicating subtle but significant differences between the cultures. In addition to understanding the nature of facial emotion recognition, this study also evaluates and enables the use of RaFD within Indian population.
Access to validated stimuli depicting children's facial expressions is useful for different research domains (e.g., developmental, cognitive or social psychology). Yet, such databases are scarce in comparison to others portraying adult models, and validation procedures are typically restricted to emotional recognition accuracy. This work presents subjective ratings for a sub-set of 283 photographs selected from the Child Affective Facial Expression set (CAFE [1]). Extending beyond the original emotion recognition accuracy norms [2], our main goal was to validate this database across eight subjective dimensions related to the model (e.g., attractiveness, familiarity) or the specific facial expression (e.g., intensity, genuineness), using a sample from a different nationality (N = 450 Portuguese participants). We also assessed emotion recognition (forced-choice task with seven options: anger, disgust, fear, happiness, sadness, surprise and neutral). Overall results show that most photographs were rated as highly clear, genuine and intense facial expressions. The models were rated as both moderately familiar and likely to belong to the in-group, obtaining high attractiveness and arousal ratings. Results also showed that, similarly to the original study, the facial expressions were accurately recognized. Normative and raw data are available as supplementary material at https://osf.io/mjqfx/.
Linguistic database project of Chinese ideophones
Experimental tasks comparing participants' performance for categorising, remembering, and recognising positive and negative words are widely used in the emotional cognitive domain. Such tasks are commonly used in experimental psychology and psychiatry research, and have been shown to be sensitive biomarkers of depression and antidepressant drug action [1,2]. In addition, several of these tasks investigate self-referential processing i.e., the processing of information relevant to oneself; this has been shown to modify the way emotional words are encoded and remembered and may be a target that is amenable to treatment [3,4]. In practice, the development of such tasks for implementation in research studies often depends on the selection and matching of words according to characteristics such as valence or arousal, imageability, word frequency and word length to investigate differences in a chosen domain of interest whilst keeping important confounds constant. This introduces a need for ratings covering a range of word attributes that have been shown to affect processing. In particular, ratings of self-referential valence (how positively or negatively subjects feel about a word when this is used to describe themselves/their circumstances) have been seldom included in databases, despite the frequent investigation of the concept in research [1,5]. Other important attributes often considered in the process of matching and selection are word imageability and subjective frequency [6,7]. To facilitate the word selection and matching process required in cognitive-emotional task development, the present dataset provides subjective ratings for 150 positive and 150 negative adjectives describing personality characteristics. Across four online surveys, the 300 words were rated on self-referential valence, imageability and subjective frequency by representative samples of 200 UK-based, English-speaking adults. Basic demographics and data on depressive symptoms and state anxiety were collected from all participants. Comprehensive descriptive statistics and word length were calculated for each of the 300 words. All data cleaning and statistical analysis was performed in R. Our work is based on years of experience using the Oxford Emotional Task Battery [1,5] and may be particularly relevant for researchers using self-referential cognitive tasks with UK-based samples.
With the Developmental Lexicon Project (DeveL), we present a large-scale study that was conducted to collect data on visual word recognition in German across the lifespan. A total of 800 children from Grades 1 to 6, as well as two groups of younger and older adults, participated in the study and completed a lexical decision and a naming task. We provide a database for 1,152 German words, comprising behavioral data from seven different stages of reading development, along with sublexical and lexical characteristics for all stimuli. The present article describes our motivation for this project, explains the methods we used to collect the data, and reports analyses on the reliability of our results. In addition, we explored developmental changes in three marker effects in psycholinguistic research: word length, word frequency, and orthographic similarity. The database is available online.
How perceptual information is encoded into language and conceptual knowledge is a debated topic in cognitive (neuro)science. We present modality norms for 643 Italian adjectives, which referred to one of the five perceptual modalities or were abstract. Overall, words were rated as mostly connected to the visual modality and least connected to the olfactory and gustatory modality. We found that words associated to visual and auditory experience were more unimodal compared to words associated to other sensory modalities. A principal components analysis highlighted a strong coupling between gustatory and olfactory information in word meaning, and the tendency of words referring to tactile experience to also include information from the visual dimension. Abstract words were found to encode only marginal perceptual information, mostly from visual and auditory experience. The modality norms were augmented with corpus-based (e.g., Zipf Frequency, Orthographic Levenshtein Distance 20) and ratings-based psycholinguistic variables (Age of Acquisition, Familiarity, Contextual Availability). Split-half correlations performed for each experimental variable and comparisons with similar databases confirmed that our norms are highly reliable. This database thus provides a new important tool for investigating the interplay between language, perception and cognition.
Este estudo apresenta dados normativos de familiaridade para utiliza{\c{c}}{\~{a}}o enquanto base para controloe manipula{\c{c}}{\~{a}}o de substantivos comuns em Portugal. Medidas de familiaridade com o referente e como significado (Larochelle {\&} Saumier, 1993) foram recolhidas em dois momentos, num primeiro casoenglobando apenas substantivos concretos ( n =320) e num segundo momento englobando tantosubstantivos concretos como abstractos ( n =219). As normas s{\~{a}}o apresentadas para um total de 459 palavras diferentes.
Words are the main components of meaning in spoken and written language. Animacy is one of the words’ attributes which is referred to the distinction between animate and inanimate stimuli, and it has an important influence on the perception of words. The main aim of the present study was to provide animacy norms for Persian words. To follow this purpose, 401 words from Warriner et al., were selected and translated into Persian. Participants (n=229) rated the words’ animacy via a 5-point Likert scale. Split-half reliability of 0.97 (p<0.001), between two halves of participants is observed.
Information on age-related differences in affective meanings of words is widely used by researchers to study emotions, word recognition, attention, memory, and text-based sentiment analysis. To date, no Chinese affective norms for older adults are available although Chinese as a spoken language has the largest population in the world. This article presents the first large-scale age-related affective norms for 2,061 four-character Chinese words (AANC). Each word in this database has rating values in the four dimensions, namely, valence, arousal, dominance, and familiarity. We found that older adults tended to perceive positive words as more arousing and less controllable and evaluate negative words as less arousing and more controllable than younger adults did. This indicates that the positivity effect is reliable for older adults who show a processing bias toward positive vs. negative words. Our AANC database supplies valuable information for researchers to study how emotional characteristics of words influence the cognitive processes and how this influence evolves with age. This age-related difference study on affective norms not only provides a tool for cognitive science, gerontology, and psychology in experimental studies but also serves as a valuable resource for affective analysis in various natural language processing applications.
Research in the socio-emotional domain may require words for experimental settings rated on emotionally and socially relevant word characteristics (e.g., valence and desirability). In addition, cognitively relevant word characteristics (e.g., imagery) are important for research in the interface of emotion and cognition (e.g., emotional memory). To provide researchers with a corresponding word pool, the database of English EMOtional TErms (EMOTE) provides subjective ratings for 1287 nouns and 985 adjectives. Nouns and adjectives were rated on valence, arousal, emotionality, concreteness, imagery, familiarity, and clarity of meaning. In addition, adjectives were rated on control, desirability, and likeableness. EMOTE norms provide an easily accessible word pool for research in the socio-emotional domain. To illustrate the usefulness of this database, norms were linked to memorability scores from a word recognition task for EMOTE nouns. The database as well as future directions are discussed.
The International Affective Picture System (IAPS) has been widely used in aging-oriented research on emotion. However, no ratings for older adults are available. The aim of the present study was to close this gap by providing ratings of valence and arousal for 504 IAPS pictures by 53 young and 53 older adults. Both age groups rated positive pictures as less arousing, resulting in a stronger linear association between valence and arousal, than has been found in previous studies. This association was even stronger in older than in young adults. Older adults perceived negative pictures as more negative and more arousing and positive pictures as more positive and less arousing than young adults did. This might indicate a dedifferentiation of emotional processing in old age. On the basis of a picture recognition task, we also report memorability scores for individual pictures and how they relate to valence and arousal ratings. Data for all the pictures are archived at www.psychonomic.org/archive/.
We describe the Age-Dependent Evaluations of German Adjectives (AGE). This database contains ratings for 200 German adjectives by young and older adults (general word-rating study) and graduate students (self-other relevance study). Words were rated on emotion-relevant (valence, arousal, and control) and memory-relevant (imagery) characteristics. In addition, adjectives were evaluated for self-relevance (Does this attribute describe you?), age relevance (Is this attribute typical for young or for older adults?), and self-other relevance (Is this attribute more relevant for the possessor or for other persons?). These ratings are included in the AGE database as a resource tool for experiments on word material. Our comparisons of young and older adults' evaluations revealed similarities but also significant mean-level differences for a large number of adjectives, especially on the valence dimension. This highlights the importance of age in the perception of emotional words. Data for all the words are archived at www.psychonomic.org/archive/.
Ecology and Technology are two keywords of the era we inhabit. Knowing how people represent these domains is essential to inform adequate interventions aimed at promoting conscious behaviors. Here we investigated this aspect by taking insights from the literature on conceptual organization. Specifically, we hypothesized Ecological and Technological concepts might have a “hybrid” nature, at the edge between Abstract and Concrete concepts. We asked a sample of Italian participants to rate 200 concepts pertaining to Ecological (e.g., deforestation), Technological (e.g., Internet), Natural (e.g., water), and Geographical/Geopolitical domains (e.g., mountain, city) on 39 semantic dimensions, some of which traditionally investigated (e.g., Context Availability), and others completely new (e.g., Political Relevance). Results indicate that Ecological and Technological concepts, despite having concrete referents, were more similar to Abstract than Concrete concepts in Concreteness~Abstractness and other semantic dimensions (e.g., Interoception, Social Valence). Interestingly, for some dimensions, they displayed a “more abstract” pattern than that of more typical Abstract concepts—e.g., later and more linguistic acquisition, higher need of others to be understood. Moreover, a Principal Component Analysis revealed three major components that explained overall the conceptual organization of our set of concepts. The first component complements the rating results, with the opposition between concreteness~abstractness, where Ecological and Technological concepts lie in the most abstract extreme. A further Hierarchical Cluster Analysis supported this distinction. Overall, our results have a twofold relevance. On a theoretical side, they contribute to enrich theories on concepts, suggesting Ecological and Technological concepts are special conceptual domains questioning the concrete-abstract dichotomy; on a more pragmatic side, they might inform societal politics on these timely themes.
Frequency, familiarity, and age of acquisition are important factors for word recognition that must be considered by researchers of language acquisition. Current psycholinguistic databases, based on studies of native English speakers, include objective frequency count, subjective rating of familiarity, and age of acquisition. One can easily employ those databases to obtain a stimulus list for one’s studies. For word recognition researchers interested in non-native English speakers in of Taiwan, however, there is currently no existing database. In this study, we created a psycholinguistic database which includes subjective familiarity rating and age of acquisition for 3,080 English words. Participants were 120 college students in Taiwan. They were asked to make judgments about 4,000 stimulus words. For recognized stimulus words, participants gave a rating of familiarity and self-report of age of acquisition; for non-recognized words, participants were asked to move on to the next stimulus word. Further analysis of the database showed that familiarity index, age of acquisition, and number of syllables are important factors for the recognition of a word. Variance in word recognition ratings for each factor was explained and implications were discussed.
A complete database is presented of relative familiarity ratings for 48 sets of Japanese words, each set comprising words overlapping in the initial portions. These ratings are of use for the generation of sets of materials for research in the recognition of spoken words.
Contains one file (abiq.dat), a database of words in French and English. This file was taken from a 3.5-inch microdisk labeled to indicate that it has accompanying tapes numbered 11, 12, and 13. The tapes were not with the microdisk, and there is no indication of who created the lexical database or tapes.
This paper presents the first large-scale corpus of French Belgian Sign Language (LSFB) available via an open access website (www.corpus-lsfb.be). Visitors can search within the data and the metadata. Various tools allow the users to find sign language video clips by searching through the annotations and the lexical database, and to filter the data by signer, by region, by task or by keyword. The website includes a lexicon linked to an online LSFB dictionary.
The Open Library for Affective Videos (OpenLAV) is a video database created for experimental emotion induction. The 188 videos of the database have a CC-BY license and were tested in a crowdsourcing study on Amazon MTurk. Valence/arousal ratings, several appraisal ratings, and emotion labels for the videos were assessed from 422 US-American participants, with an average of 71 ratings per video. Furthermore, multiple personality traits from the raters were assessed.
A database has been generated which constitutes affective ratings for single Chinese characters, viz. the Affective Norms for Chinese Characters (ANCC). This database enables researchers who study the lexico-semantic properties of Chinese characters to necessitate affective and other psycholinguistic properties in a manner that is independent and unbiased. Previous ratings for single Chinese characters, although extensive, have omitted affective properties such as valence and arousal. These factors are known to significantly influence the performance of participants in word recognition and production tasks. Close examination of the data in this study shows that affective and other psycholinguistic aspects of meaning explain a significant portion of the outcomes of previous experiments conducted on Chinese characters. The database can be accessed via OSF (osf.io/jn538).
Cite the source of the dataset as: Tjuka, Annika, Robert Forkel, and Johann-Mattis List. 2022. Linking Norms, Ratings, and Relations of Words and Concepts Across Multiple Language Varieties. Behavior Research Methods 54. 864–884. DOI: 10.3758/s13428-021-01650-1
The Open Library for Affective Videos (OpenLAV) is a video database created for experimental emotion induction. The 188 videos of the database have a CC-BY license and were tested in a crowdsourcing study on Amazon MTurk. Valence/arousal ratings, several appraisal ratings, and emotion labels for the videos were assessed from 422 US-American participants, with an average of 71 ratings per video. Furthermore, multiple personality traits from the raters were assessed.
This work introduces TCBLex, a lexical database of Finnish literary works read by children between the ages of 7 and 15. We explain in detail the work done to build the corpus TCBLex is based on, including how books were sampled and collected, turned into text files, and finally processed. We also touch on legal considerations and how it is possible to build such a corpus in the EU. TCBLex contains over 11 million tokens that are annotated with parts-of-speech tags and lemmatized. We provide 14 different sub-lexicons in total, covering individual intended reading ages, age groups, as well as different genres. We also provide versions with additional morphological information, such as the cases and tenses of words. TCBLex provides various psycholinguistically interesting lexical statistics for both word types and lemmas, such as different frequency metrics, distributions, word lengths, numbers of syllables, morphological paradigm sizes, and for the first time in a Finnish lexicon, ages when words and lemmas are first encountered in books. TCBLex is freely available at https://doi.org/10.5281/zenodo.15655580.
This repository contains the data, materials, and documentation associated with **IconicITA**, the first large scale dataset of iconicity ratings for Italian words. The resource provides iconicity norms for 1,121 words from the Italian adaptation of the ANEW (Affective Norms for English Words) database and includes ratings collected from both native Italian speakers (L1) and highly proficient Italian L2 speakers. The project investigates iconicity, namely the degree to which a word's form resembles aspects of its meaning, and its relationship with psycholinguistic variables such as perceptual strength, specificity, concreteness, age of acquisition, and word frequency. By comparing L1 and L2 judgments, the study also explores the extent to which iconicity ratings reflect language specific form meaning correspondences rather than purely semantic information. The repository supports the publication: **de Varda, A. G., Lamarra, T., Ravelli, A. A., Saponaro, C., Giustolisi, B., & Bolognesi, M. (2025). *IconicITA: Iconicity ratings of the Italian affective lexicon*. PLOS ONE, 20(12), e0337947.** The materials made available here are intended to facilitate transparency, reproducibility, and future research on iconicity, lexical semantics, psycholinguistic norms, language processing, and language mediated abstraction.
Language-and culture-specific norms are needed for research on emotion-laden stimuli. We present valence and arousal ratings for 420 Finnish nouns for a sample of 996 Finnish speakers. Ratings are provided both for the whole sample and for subgroups divided by age and gender in light of previous research suggesting age- and gender-specific reactivity to the emotional content in stimuli. Moreover, corpus-based frequency values and word length are provided as objective psycholinguistic measures of the nouns. The relationship between valence and arousal mainly showed the curvilinear relationship reported in previous studies. Age and gender effects on valence and arousal ratings were statistically significant but weak. The inherent affective properties of the words in terms of mean valence and arousal ratings explained more of the variance in the ratings. In all, the findings suggest that language- and culture-related factors influence the way affective properties of words are rated to a greater degree than demographic factors. This database will provide researchers with normative data for Finnish emotion-laden and emotionally neutral words. The normative database is available in Database S1.