1396 norm sets
Understanding the lexical characteristics of Chinese characters is crucial given their extensive usage and unique logographic structure. In this study, we normed affective ratings (valence and arousal) for 3971 Chinese characters. We investigated the relationships between intensity (mean rating) and ambiguity (rating variability) of these affective variables, alongside additional lexico-semantic variables from Su et al., Behavior Research Methods, 55(6), 2989-3008, (2022). Drawing on lexical data from 25,281 two-character words available in the Chinese Lexicon Project (Tse et al., Behavior Research Methods, 49(4), 1503-1519, 2017, Behavior Research Methods, 55(8), 4382-4402, 2023; Chan & Tse, Behavior Research Methods, 56(7), 7574-7601, 2024), we further explored cross-level relationships between character-level and word-level variables. Multiple regression analyses controlling for various lexical variables revealed several noteworthy patterns. First, we identified a quadratic valence-arousal relationship, such that characters with extreme valence ratings (either highly positive or highly negative) elicited higher arousal compared to neutral characters. This relationship was moderated by arousal ambiguity, partially consistent with previous findings (Brainerd et al. Journal of Experimental Psychology: General, 150(8), 1476-1499, 2021a), Second, we observed consistent quadratic intensity-ambiguity relationships across all variables, supporting the quadratic law proposed by Brainerd et al. Journal of Memory and Language, 121, 104286, (2021b). Finally, significant positive associations occurred between character-level variables and their corresponding word-level variables for both the first and second characters. The strength of these cross-level relationships varied across affective and lexico-semantic variables and may further be influenced by semantic transparency. Overall, our findings advance the understanding of affective and semantic features of Chinese characters and offer insights into the cross-level integration of characters' and words' lexical characteristics. The data reported in this paper are available at: https://osf.io/kh4yx.
Abstraction enables us to categorize experience, learn new information, and form judgments. Language arguably plays a crucial role in abstraction, providing us with words that vary in specificity (e.g., highly generic: tool vs. highly specific: muffler). Yet, human-generated ratings of word specificity are virtually absent. We hereby present a dataset of specificity ratings collected from Italian native speakers on a set of around 1K Italian words, using the Best-Worst Scaling method. Through a series of correlation studies, we show that human-generated specificity ratings have low correlation coefficients with specificity metrics extracted automatically from WordNet, suggesting that WordNet does not reflect the hierarchical relations of category inclusion present in the speakers' minds. Moreover, our ratings show low correlations with concreteness ratings, suggesting that the variables Specificity and Concreteness capture two separate aspects involved in abstraction and that specificity may need to be controlled for when investigating conceptual concreteness. Finally, through a series of regression studies we show that specificity explains a unique amount of variance in decision latencies (lexical decision task), suggesting that this variable has theoretical value. The results are discussed in relation to the concept and investigation of abstraction.
This study examined how well large language models (LLMs) approximate human psychological ratings for early-acquired English words. We used four state-of-the-art LLMs, including GPT-4o and Meta-Llama-3.1, to evaluate 21 static psychological features for 695 words and compared these estimates with human norms. The results showed that LLMs aligned well with human ratings for some features (e.g., Concreteness, Bodily Interactiveness) in terms of rank correlations (rs >.82) and distributional similarities but diverged notably for others (e.g., Iconicity, Arousal; rs <.48). Compared with content words, function words showed more pronounced discrepancies between human and LLM ratings. We also assessed how similarly human- and LLM-derived psychological features predicted words' age of acquisition (AoA), revealing both strong correspondences and systematic biases, depending on the model (differences in correlations ranged from -.27 to.28). Based on these analyses, we identified which features may be reliably estimated using LLMs, which require further refinement, and what methodological considerations are necessary for applying LLM-based measures in cognitive science. We discuss the implications of using LLMs as methodological tools in psychology and cognitive science, highlighting both their practical advantages (e.g., data coverage and data collection efficiency) and theoretical relevance. The present study provides a novel framework for evaluating the cognitive plausibility of LLMs by using lexical psychological features, complementing existing benchmarks.
This paper presents research on word familiarity rate estimation using the 'Word List by Semantic Principles'. We collected rating information on 96,557 words in the 'Word List by Semantic Principles' via Yahoo! crowdsourcing. We asked 3,392 subject participants to use their introspection to rate the familiarity of words based on the five perspectives of 'KNOW', 'WRITE', 'READ', 'SPEAK', and 'LISTEN', and each word was rated by at least 16 subject participants. We used Bayesian linear mixed models to estimate the word familiarity rates. We also explored the ratings with the semantic labels used in the 'Word List by Semantic Principles'.
English: This dataset provides cross-cultural name agreement norms for a set of 237 standardized color photographs. The stimuli were selected from the Bank of Standardized Stimuli (BOSS; Brodeur et al., 2010, 2014) or obtained under a Creative Commons license. We gratefully acknowledge Mathieu Brodeur, Ph.D., for authorizing the use and sharing of the BOSS photographs included in the present dataset, for which a CC-BY licence was obtained. Naming data were collected through online picture-naming questionnaires completed by adult French speakers from two linguistic and cultural contexts: Quebec French speakers in Quebec, Canada, and France French speakers in France. The dataset includes item-level norms for each French variety, including the modal name, modal name agreement, H-value, number of distinct names, two alternative names with their corresponding agreement values, and the percentage and count of blank responses. Available lexical frequency measures are also provided: objective frequency norms for France French using FreqFilms2 from Lexique 2 (New et al., 2004), and subjective frequency ratings for Quebec French using FREQ_MEAN from Desrochers and Thompson (2009). The dataset further includes the 237 photographic stimuli, a supplementary file documenting the post-collection data processing procedure, and R scripts used to compute the norms and perform statistical analyses. Participants did not consent to the deposition of the raw and processed individual-level data, or the data-processing log, in an institutional repository. The dataset is intended to support experimental, clinical, and cross-cultural research involving picture naming and culturally adapted assessment tools for French-speaking populations. -------- Français: Ce jeu de données fournit des normes transculturelles d’accord sur le nom pour un ensemble de 237 photographies couleur standardisées. Les stimuli ont été sélectionnés à partir de la Bank of Standardized Stimuli (BOSS; Brodeur et al., 2010, 2014) ou obtenus sous licence Creative Commons. Nous remercions sincèrement Mathieu Brodeur, Ph. D., d’avoir autorisé l’utilisation et le partage des photographies de la BOSS incluses dans le présent jeu de données, pour lesquelles une licence CC BY a été obtenue. Les données de dénomination ont été recueillies au moyen de questionnaires de dénomination d’images en ligne remplis par des locuteurs adultes du français issus de deux contextes linguistiques et culturels: des locuteurs du français québécois au Québec, Canada, et des locuteurs du français de France en France. Le jeu de données comprend des normes par item pour chaque variété de français, incluant le nom modal, le pourcentage d’accord sur le nom modal, la valeur H, le nombre de noms distincts, deux noms alternatifs accompagnés de leurs valeurs d’accord correspondantes, ainsi que le pourcentage et le nombre de réponses vides. Les mesures de fréquence lexicale disponibles sont également fournies: des normes de fréquence objective pour le français de France, à partir de FreqFilms2 dans Lexique 2 (New et al., 2004), ainsi que des jugements de fréquence subjective pour le français québécois, à partir de la variable FREQ_MEAN de Desrochers et Thompson (2009). Le jeu de données comprend également les 237 stimuli photographiques, un fichier supplémentaire documentant la procédure de traitement des données après la collecte, ainsi que les scripts R utilisés pour calculer les normes et réaliser les analyses statistiques. Les participants n’ont pas consenti à ce que les données brutes et nettoyées, ni le journal de nettoyage des données, soient déposés dans un dépôt institutionnel. Ce jeu de données vise à soutenir la recherche expérimentale, clinique et transculturelle portant sur la dénomination d’images et sur les outils d’évaluation adaptés culturellement aux populations francophones.
In psycholinguistic research, careful selection and control of stimuli are essential for gaining insights into cognitive processes. Within this field, pictures often serve as stimuli, which requires the use of image databases to investigate linguistic, mnemonic, and visual perceptual phenomena in different populations (children without disabilities, adults, elderly people, illiterate, and brain-damage patients; see Soares et al., 2018 for more detail).Although several image databases provide norms for variables such as naming agreement (the most common name assigned to a picture by individuals; Snodgrass & Vanderwart, 1980), conceptual familiarity (the frequency with which individuals encounter or think about the depicted object; Snodgrass & Vanderwart, 1980) and visual complexity (judgments regarding the number of lines, intricacies, and details in an image; Snodgrass & Vanderwart, 1980; see also Székely & Bates, 2000 for an objective measure of picture visual complexity), these norms are available in multiple languages but are often restricted to a limited set of black-and-white line drawings (less than 300). Notably, such black-and-white images have been found to elicit weaker recognition compared to colored pictures (Sanfeliu & Fernandez, 1996;Rossion & Pourtois, 2004). In recent decades, there has been an increased effort to develop colored image datasets in various languages. However, many of them consider small datasets (usually less than 500 pictures, but see Brodeur et al., 2014;and Krautz & Keuleers, 2022, for more extensive datasets) and/or use different normalization protocols that complicate the process of comparing data and planning and executing cross-linguistic experiments (Soares et al., 2018;Duñabeitia et al., 2022;and Zhong et al, 2024 for overviews).The Multilingual Picture (MultiPic) database (Duñabeitia et al., 2022) was designed to address the limitations of previous databases by providing researchers with norms for naming agreement and concept familiarity for a set of colored images (500), selected from an initial pool of 750 images (Duñabeitia et al., 2018). To date, this database spans thirty-three languages, including American English, Australian English, Basque, Belgium Dutch, British English, Cantonese, Catalan, Cypriot Greek, Czech, Finnish, French, German, Greek, Hebrew, Hungarian, Italian, Korean, Lebanese Arabic, Malay, Malaysian English, Mandarin Chinese, Netherlands Dutch, Norwegian, Polish, European Portuguese, Quebec French, Rioplatense Spanish, Russian, Serbian, Slovak, Spanish, Turkish, and Welsh. The images depict specific concepts and general knowledge items, and the same data collection and preprocessing protocols were consistently applied across all languages.Expanding MultiPic to include additional languages, dialects, and language varieties worldwide would enable researchers to investigate lesser-studied languages beyond the predominant focus on English, facilitate direct cross-linguistic comparisons, and deepen our understanding of cognitive processes that are universal versus those that are language-specific. The primary objective of the present study was to norm MultiPic in Galician, a relatively under-researched language, enabling researchers to conduct studies with it. Galician is a Western Ibero-Romance language predominantly spoken in Galicia, an autonomous community in northwestern Spain, where it holds co-official status with Spanish.Psycholinguistic studies on Galician are less prevalent than those on Spanish, Portuguese, Catalan, and Basque (see Comesaña & Sá-Leite, 2024). This paucity of research investigating the specific cognitive mechanisms involved in both the comprehension and production of Galician is likely attributable to several factors, including the language's recent standardization in the 1980s (Pintos, 2025) and the scarcity of databases that would allow for the careful selection of linguistic materials for various experiments. The development of these tools would reinvigorate research on Galician by ensuring that experimental outcomes accurately mirror core cognitive processes. Consequently, they would provide essential scientific evidence to inform public policies related to Galician. This language, which coexists with Spanish, presents a distinctive opportunity to examine psycholinguistic theories of language processing within bilingual contexts.Although we generally adhere to the same data collection and preprocessing procedures described in the MultiPic database (Duñabeitia et al., 2022), several adaptations were necessary to accurately reflect Galician's linguistic reality. These modifications accounted for the linguistic diversity across the region, shaped by the contact between Galician and Spanish. For instance, because Galician remains subordinate to Spanish in many social contexts, speakers often incorporate Spanish words or adapted terms, with variations across different regions (cf. Rei-Doval, 2025). Thus, we considered the diverse linguistic varieties and regional differences within Galician, ensuring that the dataset represents the full spectrum of language use across different areas. This approach not only respects the sociolinguistic context of Galician but also allows for a more comprehensive understanding of the cognitive processes involved in bilingual language processing. That is, it will enable researchers to examine which cognitive processes are general and which are specific to different sociolinguistic contexts. Nevertheless, the experimental method, the preprocessing protocol, and the data structure are comprehensively detailed to provide researchers with the necessary framework for adapting the MultiPic to other languages with comparable sociolinguistic contexts.In conclusion, MultiPic and the Galician MultiPic, in particular, serve as valuable tools that enable researchers to design studies in Galician and other languages, where the properties of the materials have been rigorously tested in parallel.The complete dataset, including the data file, is publicly available in the following repositories: https://figshare.com/articles/dataset/Untitled_Item/19328939 and https://osf.io/ank4g/?view_only=43367d3dd27543b0aa66dfb8e71ce1fc.2 MethodParticipants were recruited over two months through social media and local newspaper advertisements. Their participation was voluntary. A total of 88 Galician speakers were initially recruited, surpassing the median sample size in the original MultiPic project (i.e., 80; Duñabeitia et al., 2022)). Still, three were excluded for not following task instructions (e.g., responding in a language other than Galician or basing answers on familiarity with the picture instead of its name).From the remaining 85 participants (47 women, 34 men, four preferred not to disclose their sex; mean age = 42 [age range of 18 to 82], SD = 19.05), around 28% of participants were from O Grove, 8% from Santiago de Compostela, 8% from A Coruña, 6% from Vigo, and the rest from 24 different places. Even though all were speakers of Galician, in their daily lives, around 28% spoke only Galician, 32% spoke more Galician than Castilian Spanish, 18% spoke both languages equally, 16% spoke more Castilian Spanish than Galician, and 6% spoke only Castilian Spanish. More than half had a university degree.We used the 500 colored pictures from the MultiPic database representing common concrete concepts. These pictures were in PNG format with a 300 × 300 pixels resolution and were initially created by a local artist commissioned by the authors of the original study (Duñabeitia et al., 2018). The set of 500 elements depicted was the same as those used in Duñabeitia et al. (2022), consisting of a pictorial set of digital line drawings derived from a list of imageable and concrete Spanish words taken from ESPAL (Duchon et al., 2013).The Galician MultiPic norming followed the standardized protocol of the original MultiPic project. Instructions were provided in Galician. Sociolinguistic data, including age, gender, number of languages spoken fluently, and possession of a university degree, were collected. However, unlike other languages, the sociolinguistic data for Galician was gathered in greater detail to ensure an accurate understanding of the sociolinguistic reality of the Galician language. Thus, questions were added regarding place of birth and place of residence, age of acquisition of Galician and Spanish, language balance, educational level, and socioeconomic status.Participants received a link and completed the tasks on their computers, tablets, or phones. However, fifteen elderly adults who were not computer literate gave their responses orally, which a team member transcribed. All participants completed the two tasks in the same order using the Gorilla Experiment Builder. First, participants were provided with a link and completed the tasks by typing their responses using a computer, tablet, or smartphone. However, fifteen elderly adults who were not computer literate provided their responses orally, which a team member transcribed. Considering this, all participants named each of the 500 randomly presented images, using no more than one word per concept. Then, they rated their familiarity with each concept on a 100-point scale, ranging from 0 (not familiar at all) to 100 (very familiar). If they did not know the name of an image, they could select the "?" button, which was recorded as an "I don't know" response. Before starting, participants completed two practice trials to familiarize themselves with the procedure. The experiment lasted approximately one hour, with breaks every 50 trials. Responses were coded to account for linguistic variations, including standard (i.e., the form accepted by the Real Academia Galega [Royal Galician Academy]) versus colloquial forms, dialectal differences, and influences from Spanish. To this end, and following preceding studies (Duñabeitia et al., 2018;2022), a native speaker of Galician reviewed and corrected spelling errors while also standardizing responses by merging basic variants of the same names (e.g., hyphenated or pluralized forms).A version of the Galician MultiPic is also available at https://figshare.com/articles/dataset/Untitled_Item/19328939. However, only the data considering the Galician nouns are provided here, even when the Galician noun was not the modal name (which occurred with 52 nouns [highlighted in yellow in the dataset provided in OSF]). Thus, the Galician MultiPic at Figshare includes nine columns (from A to I) as occurs with the other 33 languages of MultiPic, which corresponds to the Language provided, the Code (number of the picture), the Number of Responses, the H statistic, the Modal Response, the Modal Response Percentage, the "I don't know" Response Percentage, the Idiosyncratic Response Percentage, and the Familiarity.Sheet H-STATISTIC contains the calculation of the H-STATISTIC for each picture.Sheet CODEBOOK contains a detailed description of the information collected in both the raw and cleaned data frames used for analyses.Regarding the naming task, two measures were considered as in earlier Multipic studies: the mean H statistic and the mean modal response percentagefoot_0. These were analyzed, and the familiarity measures were recorded as well.The most notable finding is that the data exhibits an averaged H statistic of 0.71 and a mean modal response percentage of 73.56%. As mentioned above, only 52 pictures out of 500 had one unique response. The H statistic for Galician is higher than the average for MultiPic across 33 other languages (0.55). Indeed, only 5 out of 33 languages (Malay, Lebanese, Korean, Mandarin, and Cantonese) have higher H statistic values than Galician, and only 2 (Mandarin and Cantonese) lower mean modal response percentages (73.28% and 59.17%, respectively). Interestingly, of the official languages in the Iberian Peninsula, including Basque, Catalan, Galician, Portuguese, and Castilian Spanish, Galician exhibits the highest H statistic and the lowest mean modal response percentage. In comparison, the H statistic and the mean modal response percentage for Basque are 0.66 and 82.94%, for Catalan 0.45 and 88.98%, for Portuguese 0.37 and 90.38%, and for Castilian Spanish 0.30 and 93%, respectively. Table 1 summarizes the norms for the 500 images of the MultiPic in each of the languages of the Iberian Peninsula. The relatively high mean H statistic and low mean modal response percentages obtained in the current dataset suggest a higher lexical variability when compared to most languages included in the Multipic database and to the languages that coexist in the Iberian Peninsula.Correlation analyses on the H statistic and Familiarity values across languages of the Iberian Peninsula were conducted to validate individual dataset quality. We focus on comparative analyses in these languages because speakers share not only historical and linguistic connections but also cultural and educational influences that shape familiarity judgments. This is particularly relevant for Romance languages like Castilian Spanish, European Portuguese, Catalan, and Galician, which have significant lexical and structural similarities, as well as for Basque, which, despite being non-Romance coexists in the same sociolinguistic environment. While cross-linguistic correlations can occur even between typologically distant languages, as shown in previous studies, our focus here is on a more controlled linguistic and cultural space, allowing for a more precise interpretation of familiarity effects.A correlation analysis performed on the H statistic showed that all the Pearson pairwise correlation coefficients were significant at the p < 0.001 level, with r values ranging between 0.27 (Galician vs. European Portuguese) and 0.59 (Spanish vs. Catalan). The reason why r values between Galician and European Portuguese are the lowest despite their status as closely related languages with a shared medieval history as part of Galician-Portuguese, may be attributed to the distinct sociolinguistic contexts in which they have developed. These differing contexts have played a significant role in shaping lexical variation between the two languages. Note that Galician coexists with Castilian Spanish, a language of high prestige, which has led to the incorporation of numerous lexical borrowings from this language into Galician (see Dubert, 2025). Furthermore, the establishment of an official written standard for Galician did not occur until 1980, highlighting the relatively recent process of linguistic standardization. In contrast, Portuguese is the main language in Portugal and does not coexist with another widely spoken language, except for Mirandese, which is used in the specific region of Miranda do Douro. Additionally, Portuguese has a long-established linguistic tradition, with its first grammar and dictionary dating back to the 16th century. Since these early efforts, a strong normative tradition has been maintained (Santos, 2018).Likewise, the correlation analyses performed on different familiarity scores obtained for each item in each language showed that all the Pearson pairwise correlation coefficients were significant at the p < 0.001 level, with r values ranging between 0.79 (Catalan vs. Basque) and 0.83 (Galician vs. European Portuguese).Besides, a correlation analysis was conducted between the H statistic and Familiarity values in each official or co-official language from the Iberian Peninsula already tested in the MultiPic database. We found low to moderate negative and significant correlations in all of them. That is, the higher the values in familiarity, the lower the values in the H statistic, which makes sense as the higher the H statistic, the lower the name agreement. To be more precise, all Pearson pairwise correlation coefficients were significant at the p < 0.001 level, except for the Galician language ( p =.02), with r values ranging between -.10 (for the Galician) and -.44 (for the Catalan). The smallest correlation was found for Galician. At first, we thought that this was probably because it has greater lexical variability than the other languages. Indeed, if we look at the second language from the Iberian Peninsula that has a high lexical variability (Basque), we can see that it also showed a small correlation value between H statistic and Familiarity (-.25). However, when compared with other languages like Chinese or Malay that also have a great lexical variability we found high significant correlations (-.49 and -.90, respectively). Therefore, a more plausible explanation may lay on the fact that familiarity modulates agreement (and not the other way around). That is, if someone is not familiar (or that much familiar) with an object, they would be hesitant when naming it, and as a consequence, this would lead to lower agreement scores across participants. This would be true for all the languages. However, for Galician more variables than familiarity may be explaining this result such as the already mentioned coexistence with the Castilian language, the recent official written standard for Galician, which means that it is not perfectly implemented, and the desire of some people to reflect their dialectal variant. We recognize, however, that this is a tentative explanation that deserves further examination.Each variable's inter-rater reliability was determined by calculating intraclass correlations (ICCs) via a two-way random consistency model. ICCs revealed acceptable reliability for H statistic (ICC = 0.78 [0.75, 0.81]), and an excellent reliability for familiarity (ICC = 0.92 [0.91, 0.93]).Although all correlations were significant, findings underscore how social and regional factors, such as dialectal variation, hyper-Galician forms, and Spanish influence, shape lexical variation in Galician. A closer look at the responses given to each picture shows that some pictures had multiple interpretations that seem to reflect the variety of realities of the population, i.e., participants used nouns with different meanings (e.g., magnet vs. horseshoe), for example, by using different co-hyponyms or elements of the same semantic field ( figo vs. cebola [fig vs. onion]), or, in some cases, by focusing on different elements or areas of the image (e.g., for the picture of a shoulder, participants used names like costas, pel, or marrón, i.e., back, skin, or brown in English). On the other hand, some pictures were named with synonyms, such as xornal and periódico, two different Galician words to name a newspaper. Importantly, in many cases, the noun participants used depended on the dialectal variety of their region (e.g., vespa, avespa, and avéspora for wasp). Also, in some cases, participants used "hyper-Galician" forms, i.e., linguistic that are created when speakers to use they as or Galician. often these to the of Castilian Spanish, a shaped by the unique context of language contact and the relatively standardization of Galician in Galicia, as For example, the Galician form for is, but the hyper-Galician form is many of lexical from Spanish can be such as the Spanish word for or for In in more than of the images the modal form is the Spanish word (e.g., or or the adapted Spanish word (e.g., and in and Spanish, or and in and Spanish, may however, that the lexical variability in Galician is by the pictures than by the of the language as an In other the same concepts consistently the lowest agreement across languages. This does not seem to be the when calculating the mean H statistic of the Multipic database including all the languages tested this was = and the mean modal response percentage was = These values are to those provided in earlier studies with different of stimuli (e.g., et al., et al., et al., et al., et al., et al., and as Duñabeitia et al. have already out when comparing the data of languages, relatively low mean H statistic and the high mean modal response percentages of the current dataset suggest high name agreement across items, languages, and the materials for their use in different of experiments and these The of data from specific regions and the written of the experiment questions about and to the relatively recent standardization of the Galician language, many speakers are with the spelling or This questions about the of this task in a written Nevertheless, this the of comprehensive planning and cultural when the Galician MultiPic valuable insights for cross-linguistic studies and research, recognition of linguistic diversity while providing a framework for adaptations in other Galician of MultiPic a in psycholinguistic research by providing standardized norms for an language. This the lexical variation in Galician, shaped by regional and language contact with like regional and Multipic is a written than spoken the Galician MultiPic is an essential for cross-linguistic studies and the of bilingual cognitive valuable insights for research on language
The publication of the Vai dictionary based on August Klingenheben’s collection and edited by Raimund Kastenholz is a very fortunate event for African studies. It finally makes the rich lexical database that was compiled by August Klingenheben (1886-1967) accessible. The thorough and consistent editorial work carried out by Kastenholz has made available the results of Klingenheben’s research and documentation project on Vai, which, as we learn from the introduction, spanned a period of over f...
This paper introduces and describes the second edition of the Mixtec Sound Change Database (Auderset & Campbell, 2024), which now includes a module on tone change. Tone change is an under-researched topic in historical linguistics and is virtually absent from cross-linguistic databases. The Mixtec languages from southern Mexico provide an ideal starting point for studies in this area, as they are all tonal and tone can be reconstructed to the proto-language. This updated version of the database thus serves as a model for including tonal data in synchronic and diachronic studies of languages and language families where this feature is present. This will support efforts towards a better understanding of suprasegmental change in Mixtec and beyond.
Valence, Arousal and Dominance ratings for over 13k Italian content words (including nouns, adjectives and verbs)
Abstract In the present study, we developed affective (valence and arousal) and sensory–motor (concreteness and imageability) norms for 210 English idioms rated by native English speakers (L1) and English second-language speakers (L2). Based on internal consistency analyses, the ratings were found to be highly reliable. Furthermore, we explored various relations within the collected measures (valence, arousal, concreteness, and imageability) and between these measures and some available psycholinguistic norms (familiarity, literal plausibility, and decomposability) for the same set of idioms. The primary findings were that (i) valence and arousal showed the typical U-shape relation, for both L1 and L2 data; (ii) idioms with more negative valence were rated as more arousing; (iii) the majority of idioms were rated as either positive or negative with only 4 being rated as neutral; (iv) familiarity correlated positively with valence and arousal; (v) concreteness and imageability showed a strong positive correlation; and (vi) the ratings of L1 and L2 speakers significantly differed for arousal and concreteness, but not for valence and imageability. We discuss our interpretation of these observations with reference to the literature on figurative language processing (both single words and idioms).
Emotion lexicons are useful in research across various disciplines, but the availability of such resources remains limited for most languages. While existing emotion lexicons typically comprise words, it is a particular meaning of a word (rather than the word itself) that conveys emotion. To mitigate this issue, we present the Emotion Meanings dataset, a novel dataset of 6000 Polish word meanings. The word meanings are derived from the Polish wordnet (plWordNet), a large semantic network interlinking words by means of lexical and conceptual relations. The word meanings were manually rated for valence and arousal, along with a variety of basic emotion categories (anger, disgust, fear, sadness, anticipation, happiness, surprise, and trust). The annotations were found to be highly reliable, as demonstrated by the similarity between data collected in two independent samples: unsupervised (n = 21,317) and supervised (n = 561). Although we found the annotations to be relatively stable for female, male, younger, and older participants, we share both summary data and individual data to enable emotion research on different demographically specific subgroups. The word meanings are further accompanied by the relevant metadata, derived from open-source linguistic resources. Direct mapping to Princeton WordNet makes the dataset suitable for research on multiple languages. Altogether, this dataset provides a versatile resource that can be employed for emotion research in psychology, cognitive science, psycholinguistics, computational linguistics, and natural language processing.
SUBTLEX-SR is a subtitle-based frequency norm for Serbian, the Serbian member of the SUBTLEX family of psycholinguistic frequency resources (following Brysbaert & New, 2009). It provides word-form and lemma frequencies, contextual diversity, and dispersion measures derived from the Serbian portion of OpenSubtitles v2018. **Contents.** The resource consists of two lexical tables: - **Wordform table** (2,198,809 entries): one row per surface form, with frequency from a 50-million-token lemmatized subsample, frequency from the full 64,842-film cleaned corpus, contextual diversity, three dispersion measures (Gries DP, DPnorm, Juilland D over decade buckets), and POS distribution.- **Lemma table** (330,535 entries): one row per lemma, with subsample-derived frequency, contextual diversity, and POS distribution. Both tables are provided in two scripts (Latin and Cyrillic) and in two formats (CSV and Apache Parquet). Methodology JSON files documenting all construction decisions, and a deterministic per-film manifest of the lemmatized subsample, are included for reproducibility. Four figures from the accompanying paper (baseline correlations, register divergence, Zipf distribution, CD vs. frequency) are also included. **Headline numbers.** 64,842 distinct films (the contextual-diversity base); 287.6 million alphabetic tokens in the cleaned corpus; 50.5 million classla-tokens in the lemmatized subsample; year coverage 1902–2020 with frequency data restricted to films from 1950 onwards. **Construction summary.** Source: OpenSubtitles v2018 raw Serbian (178,596 XML files). Cleaning included repair of a CP-1250-as-CP-1252 mojibake encoding error affecting 90.6% of source documents, two-phase deduplication consolidating 178,494 cleaned uploads into 64,842 distinct films, language-script normalization, and contamination filtering. Lemmatization performed with classla 2.2.1 (standard model variant) on a stratified subsample of 8,933 films. **Validation.** SUBTLEX-SR correlates strongly with OPUS's pre-computed Serbian frequency table (Pearson *r* = 0.97 on 498,489 forms; internal-consistency check) and moderately with the web-derived srLex baseline (*r* = 0.68 on 216,260 forms; *r* = 0.69 on 35,968 lemmas). The reduced srLex correlation reflects a register difference between subtitle dialogue and web prose; the divergent vocabulary sorts coherently into dialogue-characteristic classes (negated future-tense auxiliaries, vocatives, interjections) on one side and news/government-characteristic classes (country and region names, politicians' surnames, formal connectives) on the other. **Limitations.** No behavioral validation against lexical decision RT data has been performed for this release; this is being pursued through collaboration with the Laboratory for Experimental Psychology, University of Novi Sad. classla's Serbian model lemmatizes a small set of Serbian forms to their Croatian variants (e.g., *šta → što*, *koga → tko*); this affects all classla-sr users and is documented in the README. The corpus contains translation residue from non-Serbian source films (predominantly English-language Hollywood and BBC content) and some Bosnian/Croatian-orthography forms typical of the BCMS continuum. See the README and methodology JSON files for full documentation. **Companion paper.** Popović, M. (in preparation). *SUBTLEX-SR: A subtitle-based frequency norm for Serbian, with attention to register coverage in existing Serbian frequency resources.* Submitted to Language Resources and Evaluation. **License.** CC BY-SA 4.0. Source subtitle text is not redistributed with this deposit; see the README for source-corpus access via OPUS. **References.** - Brysbaert, M., & New, B. (2009). Moving beyond Kučera and Francis: A critical evaluation of current word frequency norms and the introduction of a new and improved word frequency measure for American English. *Behavior Research Methods*, 41(4), 977–990.- Lison, P., & Tiedemann, J. (2016). OpenSubtitles2016: Extracting large parallel corpora from movie and TV subtitles. *Proceedings of LREC 2016*, 923–929.- Ljubešić, N., & Dobrovoljc, K. (2019). What does neural bring? Analysing improvements in morphosyntactic annotation and lemmatisation of Slovenian, Croatian and Serbian. *Proceedings of BSNLP 2019*, 29–34.
In this paper, we present our recent experience in constructing a first-of-its-kind functional corpus based on the theoretical framework of Systemic Functional Linguistics. Annotated on selected texts from the Penn Treebank, the corpus was built by a collaborative team on a web-based annotation platform with several advanced features. After a discussion on the background and motivation of the project, we present our solutions to some of the challenges encountered in the collaborative annotation process. With fine-grained annotations of an initial corpus now available, the corpus can serve as a valuable linguistic resource that complements existing semantically annotated corpora and aids in the development of a larger-scale resource crucial for automated systems for analysis of linguistic function.
. In Study 2, we used estimates and human ratings to predict valence effects in a lexical decision task. Again, we found that norms derived from vector space models and label lists for adults outperformed other combinations and approximated the functional form of children's and adults' valence effects best. We discuss the practical implications of these findings. Additionally, a new set of valence ratings from children (mean age = 12.5 years) for 535 German words is made available.
A long-standing goal shared by researchers has been to design optimal experimental procedures, including the selection of appropriate stimuli. Pictures are commonly used in different research fields. However, until recently, researchers have relied mostly on line-drawings, which can have poor ecological validity. We developed a set of high quality standardized photographs of objects from six different categories, recorded under two camera viewpoints, and five presentation conditions (on its own, held by clean hands, and by hands covered with different substances: sauce, chocolate and mud). These various staging conditions can be used to induce different emotional states while maintaining the object of interest constant. We first report normative data on the objects' name agreement and familiarity collected from North American and Portuguese participants. Results showed high name agreement and familiarity in both samples. Next, arousal, disgust and valence ratings were collected for the stimuli under either an emotional-activating or a neutral context. Subjective ratings varied according to the staging condition and the context, confirming that the same items can effectively be used in different emotional conditions. This database allows researchers to select more ecologically-valid stimuli according to their research purposes while considering several variables of interest and avoiding item-selection problems commonly present when comparing responses to neutral and emotional items.
Normative data for naming photographs are essential in psycholinguistic research. However, image naming norms are typically derived from young adults, limiting their relevance for older populations, who are at greater risk for language impairments due to neurological conditions such as stroke, traumatic brain injury or dementia. Further, lexical retrieval declines also in healthy aging, making it essential to establish norms for older adults to distinguish normal from impaired word retrieval. This study provides normative data for 600 photographs of the Bank of Standardized Stimuli (BOSS) focusing on three age cohorts (40-50, 51-65, and 66+). We examined naming accuracy, name agreement, H values, and response times (RT) to explore age-related differences in image naming. Participants completed a web-based oral picture naming task via video conferencing. Results revealed overall high naming accuracy (mean = 80.5%) and name agreement (mean = 87.4%) across the full sample, with modest variability across the range of adults self-reportedly free of neurological deficits. The 51-65 cohort showed the highest accuracy and fastest RTs. Significant correlations between RT and name agreement and H value support the inclusion of RT as key indices of naming difficulty. We discuss the implications of these findings considering psycholinguistic norms, demographic influences, and methodological differences from previous image norming studies. Novel contributions of this study include normative data for a large sample of middle to older age adults including RT and alternative names, expanding the utility of the BOSS image set for examining aging-related changes in lexical access. The study underscores the importance of including RT measures alongside traditional naming norms for improved characterization of visual stimuli. Open access to the updated dataset aims to facilitate future research into age-related language processing and supports personalized applications in cognitive and clinical settings.
Project description This project hosts TUNorms, a set of psycholinguistic word norms for Thai developed to support research on lexical–semantic processing and embodied cognition. The dataset provides normative ratings for imageability, body–object interaction (BOI), and subjective frequency for 627 mono- and multi-syllabic Thai words. The norms were collected from Thai university students using standardised rating procedures. Reliability was assessed through internal consistency and cross-linguistic comparisons with existing norms for overlapping items, and the measures were further validated in a semantic categorisation task demonstrating independent effects of imageability and BOI on response latencies after controlling for established lexical variables. The repository contains the aggregated normative data and the full rating instructions (Thai and English). Raw participant-level data are not included, as this is a completed normative study and the shared materials are intended to support controlled stimulus selection, replication, and secondary analyses. This dataset accompanies a journal manuscript currently under submission. Upon acceptance, the final citation will be added to this record. The norms are released under a Creative Commons Attribution 4.0 International (CC BY 4.0) licence to facilitate reuse.
We introduce a database (IDEST) of 250 short stories rated for valence, arousal, and comprehensibility in two languages. The texts, with a narrative structure telling a story in the first person and controlled for length, were originally written in six different languages (Finnish, French, German, Portuguese, Spanish, and Turkish), and rated for arousal, valence, and comprehensibility in the original language. The stories were translated into English, and the same ratings for the English translations were collected via an internet survey tool (N = 573). In addition to the rating data, we also report readability indexes for the original and English texts. The texts have been categorized into different story types based on their emotional arc. The texts score high on comprehensibility and represent a wide range of emotional valence and arousal levels. The comparative analysis of the ratings of the original texts and English translations showed that valence ratings were very similar across languages, whereas correlations between the two pairs of language versions for arousal and comprehensibility were modest. Comprehensibility ratings correlated with only some of the readability indexes. The database is published in osf.io/9tga3, and it is freely available for academic research.
A verb as the fundamental part of a sentence is important and its retrieval consists of different cognitive stages. Additionally, verb retrieval difficulty is reported in some types of aphasia and other neurological diseases, and some psycholinguistic variables can influence the verb retrieval process. This study aimed to provide a normative database in the Persian language for 92 black and white action pictures and related verbs in two groups of young ages (20 to 40 years old) and middle ages (41 to 64 years old). A total of 150 volunteers participated in this study, and the groups had similar characteristics due to education. The pictures were normed for variables such as name agreement, familiarity, visual complexity, age of acquisition, and image agreement. Correlation coefficients were calculated values among these measures, and comparisons were made between the two age groups. The results of the comparisons between the two groups showed that name agreement and familiarities were age-dependent. The results revealed that all measures varied with age. Also, the present study provided a set of verbs and their pictures in the Persian language and normative data were obtained on the psycholinguistic variables such that it can be used for clinical practice and research in the areas of verb processing and their naming.
The concreteness-abstractness continuum is considered a primary dimension in the representation of semantic networks. Its theoretical importance and clinical significance are widely acknowledged. To assist and enhance future research, this study collected and evaluated concreteness/abstractness ratings for 9,877 two-character Chinese words retrieved from the MEga study of Lexical Decision in Simplified CHinese (MELD-SCH, Tsang et al, 2018). The ratings were validated through comparisons with previous rating studies on concreteness and imageability of smaller word samples. Relations of word concreteness with word frequency, age-of-acquisition, and efficiency of lexical processing were also examined. These ratings provide an additional dimension of information to two-character words in the database MELD-SCH, permitting not only more comprehensive research on the Chinese language, but also cross-language investigation of the concreteness effect between Chinese and other languages such as English and Dutch where a large database of concreteness ratings is also available.
Word characteristics such as frequency, imageability, concreteness and length are considered good predictors of performance in lexical tasks like picture naming, word comprehension or lexical decision-making. There is also evidence that the age of acquisition (AoA) of words can partly explain aspects of word processing behaviour in later childhood and adulthood (Morrison et al., 1992; Brysbaert & Cortese, 2010).In the present study, we collected AoA norms for 158 nouns and 142 verbs in 22 languages: Afrikaans, British English, Catalan, Danish, Finnish, German, Hebrew, Irish, IsiXhosa, Italian, Lithuanian, Luxembourgish, Maltese, Norwegian, Polish, Russian, Serbian, Slovak, South African English, Spanish, Swedish and Turkish. In a preparatory picture naming procedure, adult native speakers of 34 languages were asked to name 508 object and 504 action pictures. Words shared among the target languages were retained for the final corpus. Our study followed the typical procedure for establishing AoA (see Morrison et al. 1997) and was performed on-line (see www.words-psych.org). 804 adult participants (at least 20 for each language) were asked to specify the age at which they learned the words in their native language. The vast majority of words were rated as acquired by the age of 7 years, demonstrating overlap in early vocabulary across diverse languages. Significant correlations between all language pairs point to a similar developmental sequence for the words under investigation. No previous study has compared AoA judgements on a shared set of words in a wide range of languages. 'The AoA data collected in the 22 languages provides word characteristics that should assist the design of cross-linguistic psycholinguistic experiments and the preparation of materials for use in the assessment and treatment of language disorders in preschool children. The AoA data are currently being used to control for AoA in the construction of cross-linguistic lexical tasks assessing word knowledge in monolingual and bilingual children.
This project aims to create the first open normative database of Basque emotion words by collecting emotional valence, arousal, and concreteness ratings for 4,200 words from native and highly proficient Basque speakers, using ANEW-based methods and making the results openly available for research and language technology applications. The poster was presented at the fourth CLARIAH-EUS workshop, on November 28, 2025 in Vitoria-Gasteiz.
Concreteness is a fundamental dimension of word semantic representation that has attracted more and more interest to become one of the most studied variables in the psycholinguistic and cognitive neuroscience literature in the last decade. Concreteness effects have been found at both the brain and the behavioral levels, but they may vary depending on the constraints of the context and task demands. In this study, we collected concreteness norms for English and Italian words presented in different context sentences to allow better control and manipulation of concreteness in future psycholinguistic research. First, we observed high split-half correlations and Cronbach's alpha coefficients, suggesting that our ratings were highly reliable and can be used in Italian- and English-speaking populations. Second, our data indicate that the concreteness ratings are related to the lexical density and accessibility of the sentence in both English and Italian. We also found that the concreteness of words in isolation was highly correlated with that of words in context. Finally, we analyzed differences between nouns and verbs in concreteness ratings without significant effects. Our new concreteness norms of words in context are a valuable source of information for future research in both the English and Italian language. The complete database is available on the Open Science Framework (doi: 10.17605/OSF.IO/U3PC4).
A pluricentric language is a language that is used in at least two countries where it has the official status of a state, commonwealth or regional language with at least partially its own (codified) norms that usually contribute to the personal identity of speakers. Pluricentric languages have one dominant variant and (one or) several non-dominant varieties. As a result of the political fragmentation of the Hungarian language area that developed after the First World War, and then, confirmed by the peace treaties after the Second World War, the Hungarian language is one of the pluricentric languages in Europe. The article examines the results of close linguistic contacts in non-dominant varieties of the modern Hungarian language used outside Hungary. The consequences of language contacts are highlighted on the basis of lexical borrowings, which are fixed in a specific online dictionary. The dictionary consists of borrowed words of foreign origin used by autochthonous Hungarian minorities living in the Carpathian Basin outside Hungary. In addition to words and phrases that are used exclusively in the speech and writing of Hungarians in countries neighboring Hungary, words that are also used in Hungary, but with a different meaning, were also collected in the database. As of the end of September 2022, the dictionary database contained 5,034 dictionary entries (words). Since this online loanword list contains direct borrowings from many languages of the Carpathian Basin that are in contact with Hungarian (mostly from the official or state languages of Hungary's neighboring countries, including Slovak, Ukrainian, Romanian, Serbian, Croatian, Slovenian, and German), the database is a rich source for the study of contacts between Hungarian and Indo-European languages. Based on the material of the online dictionary, it was found that among the lexical borrowings of the Hungarian language –as a result of centuries-old contacts between Hungarian and various Slavic languages –borrowings of Slavic origin constitute the largest layer of vocabulary of foreign origin in the Hungarian language. The result of the project is a dictionary database that provides an opportunity for a comparative analysis of the vocabulary of non-dominant variants of the pluricentric Hungarian language.
<ns4:p>Over the past decade, there have been several attempts to standardize cross-linguistic datasets. Since language comparison is a notoriously difficult endeavor, it requires tools that facilitate standardization and are convenient to use. The Concepticon is based on a toolkit provided for cross-linguistic comparison and offers a reference catalog for comparable concepts that appear in concept lists. While curating the Concepticon, we found that a variety of studies in distinct research fields collected information on word properties. However, until recently, no resource existed that contained these data to enable the comparison of the different word properties across languages. This gap was filled by the Database of Norms, Ratings, and Relations (NoRaRe), which is an extension of the Concepticon. Here, we present the major release of both resources - Concepticon Version 3.0 and NoRaRe Version 1.0 - which represents an important step in our data development. We show that extending and adapting the data curation workflow in Concepticon to NoRaRe is useful for the standardization of cross-linguistic datasets. In addition, combining datasets from different research fields enables studies grounded in language comparison. Concepticon and NoRaRe include lexical data for various languages, tools for test-driven data curation, and the possibility for data reuse. The first major release of NoRaRe is also accompanied by a new web application that allows convenient access to the data.</ns4:p>
Anxiety is the unease about a possible future negative outcome. In recent years, there has been growing interest in understanding how anxiety relates to our health, well-being, body, mind, and behaviour. This includes work on lexical resources for word-anxiety association. However, there is very little anxiety-related work on larger units of text such as multiword expressions (MWE). Here, we introduce the first large-scale lexicon capturing descriptive norms of anxiety associations for more than 20k English MWEs. We show that the anxiety associations are highly reliable. We use the lexicon to study prevalence of different types of anxiety- and calmness-associated MWEs; and how that varies across two-, three-, and four-word sequences. We also study the extent to which the anxiety association of MWEs is compositional (due to its constituent words). The lexicon enables a wide variety of anxiety-related research in psychology, NLP, public health, and social sciences. The lexicon is freely available: https://saifmohammad.com/worrylex.html
Naturalistic paradigms provide ecologically valid insights into affective and cognitive processes but often require costly and time-consuming human annotations. Large language models (LLMs) provide a scalable tool for generating human-like affective ratings that could complement traditional behavioral approaches. In this study, we compared affective ratings of narrative segments obtained from young adults, five OpenAI's GPT models, Meta's Llama 3.1, and lexical-level norms from the SCOPE metabase. LLM-derived ratings of hedonic valence showed strong correlations with human ratings and outperformed lexical norms. When applied to fMRI data, LLM-derived ratings identified affective brain networks that substantially overlapped with those revealed by human ratings. These findings demonstrate that LLMs can approximate group-level affective ratings from young adults in naturalistic contexts and serve as a useful complement to traditional behavioral data collection, while underscoring the need for careful evaluation of their generalizability and potential biases.
This article presents AI-generated estimates for five characteristics of German words: concreteness, valence, arousal, age of acquisition (AoA), and word familiarity. The estimates were generated using GPT-4o-mini, which was selected due to its good performance in previous studies. Validation studies were conducted comparing the AI-generated estimates with both human ratings and previously generated AI data to ensure their usefulness for research applications. The main results are as follows. The GPT estimates of word concreteness, valence, and arousal show a strong correlation with human ratings but are not better than the best available AI-generated estimates based on semantic vectors. The GPT estimates of AoA are good approximations of human ratings and outperform other available alternatives (except for human ratings), especially after the model was fine-tuned based on 2,000 human ratings. Fine-tuned AI-generated estimates of word familiarity have better predictive value than word frequency for word recognition in lexical decision tasks and vocabulary tests. Estimates for concreteness, valence, arousal, and AoA are available for 167,000 words, which are likely to be known to more than 90% of participants in typical adult studies. Word familiarity estimates are presented for 928,000 word forms. All data and codes, including newly collected human familiarity ratings for 11,000 words, are publicly available at https://osf.io/ghjd2/. The data may be freely used for research purposes, but not for commercial purposes.
This paper presents a new corpus of 140 high quality colour images belonging to 14 subcategories and covering a range of naming difficulty. One hundred and six Spanish speakers named the items and provided data for several psycholinguistic variables: age of acquisition, familiarity, manipulability, name agreement, typicality and visual complexity. Furthermore, we also present lexical frequency data derived internet search hits. Apart from the large number of variables evaluated, these stimuli present an important advantage with respect to other comparable image corpora in so far as naming performance in healthy individuals is less prone to ceiling effect problems. Reliability and validity indexes showed that our items display similar psycholinguistic characteristics to those of other corpora. In sum, this set of ecologically valid stimuli provides a useful tool for scientists engaged in cognitive and neuroscience-based research.
Normed linguistic stimuli are fundamental in psycholinguistics because they capture lexical and semantic properties that influence comprehension. However, generating these norms at scale is challenging, often leading researchers to rely on ad hoc norms collected from small samples, which can introduce inconsistencies and limit cross-study comparisons. In the present study, we investigated how large language models (LLMs) can support psycholinguistic research by prompting eight current LLMs to norm 300 English two-word metaphor combinations, such as sharp mind. We selected the dimensions of familiarity, aptness, concreteness, metaphoricity, and constituency, as these tap distinct cognitive processes and may provide insight into which aspects LLMs capture accurately and which they do not. We varied stimulus presentation (in context vs. in isolation) and response format (categorical vs. numerical) to examine which manipulation yields norms most closely aligned with human ratings. We then assessed the reliability and validity of model responses and used them to replicate existing analyses of metaphor comprehension. Overall, LLM-generated norms aligned best with familiarity and metaphoricity, which rely on word co-occurrence. In contrast, aptness, concreteness, and constituency—which require reasoning about the relationship between the topic (e.g., mind) and the vehicle (e.g., sharp)—proved more challenging for LLMs.
Contains all data from this psycholinguistic study, including original surveys, compiled data, and analysis R code. Part of the toolkit of language researchers is formed of stimuli that have been rated on various dimensions. The current study presents modality exclusivity norms for 336 properties and 411 concepts in Dutch. Forty-two respondents rated the auditory, haptic, and visual strength of these words. Mean scores were then computed, yielding acceptable reliability values. Measures of modality exclusivity and perceptual strength were also computed. Furthermore, the data includes psycholinguistic variables from other corpora, covering length (e.g., number of phonemes), frequency (e.g., contextual diversity), and distinctiveness (e.g., number of orthographic neighbours), along with concreteness and age of acquisition. To test these norms, Lynott and Connell’s (2009, 2013) analyses were replicated. First, unimodal, bimodal, and tri-modal words were found. Vision was the most prevalent modality. Vision and touch were relatively related, leaving a more independent auditory modality. Properties were more strongly perceptual than concepts. Last, sound symbolism was investigated using regression, which revealed that auditory strength predicted lexical properties of the words better than the other modalities did, or else with a different direction. All the data and analysis code, including a web application, are available from https://osf.io/brkjw/. Data and analyses dashboard: https://pablobernabeu.shinyapps.io/Dutch-modality-exclusivity-norms/ (in case of downtime, please visit https://pablobernabeu.github.io/dashboards/Dutch-modality-exclusivity-norms/d.html). Online RStudio environment with data and code: https://mybinder.org/v2/gh/pablobernabeu/Modality-exclusivity-norms-747-Dutch-English-replication/master?urlpath=rstudio Paper (Bernabeu, 2018): https://psyarxiv.com/s2c5h The norms were used, and validated, in an experiment that implemented the conceptual modality switch. Data for that experiment may be found as a linked component of the present Project.
The age of acquisition (AoA) refers to the age at which an individual learns specific items or words. Research on word recognition has shown that items with lower AoA—those acquired earlier—can be processed more quickly and accurately. In the field of Japanese word recognition, a large-scale database that would enable mega-study approaches has not been well-established. In this study, we developed an AoA norm for over 5,000 Japanese words. A total of 1,345 adults rated the AoA of 5,736 words using a 7-point scale. These ratings demonstrated satisfactory reliability. Furthermore, when examining the correlation with lexical decision task performance, words with lower AoA showed shorter response times and higher accuracy than words with higher AoA. These findings indicate that the AoA rating database collected in this study serves as a valuable resource for research using Japanese words.
Auditory pseudowords are widely used in psycholinguistics and cognitive neuroscience, but their construction requires control of sublexical familiarity and careful characterization of how acoustic cue manipulations may shift perceived lexical plausibility. Here we introduce the Minho Pseudoword Wordlikeness Ratings (MPWR), the first normative dataset of wordlikeness judgments for European Portuguese (EP) auditory trisyllabic CV pseudowords, and evaluate whether adding a localized F0-based prominence cue modulates wordlikeness beyond distributional familiarity. One hundred and twenty pseudowords were assembled from naturally produced syllables drawn from the Minho Spoken Syllable Pool (MSSP) and recorded under uniform conditions. Each item was implemented in three token types with constant segmental content: a flat baseline and two F0-enhanced versions (+15%) targeting either the penultimate or final syllable. Native EP listeners (N = 101) provided wordlikeness ratings on a 7-point scale. MSSP-derived indices quantified pseudoword syllable familiarity (SWIAll, SWIN3) and stress-position propensity for the targeted syllable (SPPmarked). Ratings were intentionally low overall yet showed substantial item-to-item variability. F0 enhancement produced a small but reliable decrease in wordlikeness relative to flat tokens, with no reliable difference between penultimate and final targeting positions. SWIAll robustly predicted ratings, whereas SPPmarked added little explanatory value. MPWR provides a practical EP resource for selecting and matching auditory pseudowords using normative wordlikeness ratings and transparent corpus-based descriptors.
Abstract The functional variants of International English are often differently distributed in the different regional standards. With evidence from the corpus of Australian English, this has already been shown for lexical variants such as will/shall, maybe/perhaps etc. In this paper evidence from the Australian corpus is used to discuss a number of variables in a) morphology b) the system of conjunction c) the system of quantifiers. The redistribution of morphological variants-edl-t (as in burned/burnt), and -wards(s) (as in downward(s)) showed a tendency to assign different grammatical roles to each variant. Among the conjunctions, apart from individual differences the most interesting finding was the higher level overall in the use of subordinating conjunctions, when Australian newspaper data was compared with the equivalent in Britain or America. A possible explanation for this invokes the Hallidayan principle that subordination is actually more common in speech than in writing. The suggestion is that Australian press reporting approximates more closely to spoken than to written norms of language. But on the quantifiers a few/several the corpus provides no support for a new popular use of several, to mean vaguely large number.
Our internal repository of words, often known as the mental lexicon, has primarily been modelled by psychologists as some kind of network. One way to probe its organisation and access mechanisms is by means of word association techniques, which have rarely been applied on Chinese. This paper reports on the design and implementation of a pilot word association test on native Hong Kong Cantonese speakers. The test contains 500 stimulus words, carefully selected and controlled on important factors including word frequency, part-of-speech, syllabicity, concreteness and vocabulary type. The resulting association norms based on 58 participants reveal interesting properties of the Chinese mental lexicon, such as the dominance of disyllabic and nominal concepts, and collocational associations. Despite its current small scale, the word association norms obtained from this study do not only offer first-hand psycholinguistic evidence for investigating the Chinese mental lexicon but also provide a useful resource to inform future studies in Chinese lexical access, lexical semantics and lexicography. 1
A total of 1,363 images from seven sets of facial stimuli were normed using the self-assessment manikin procedure. Each participant provided valence, arousal, and dominance ratings for 120-130 faces displaying various emotional expressions (e.g., happiness, sadness). The current work provides a large database of normed ratings for facial stimuli that complements the existing International Affective Picture System and the Affective Norms for English Words that were developed to provide a normative set of emotional ratings for photographs and words, respectively. This new database will increase experimental control in studies examining the perception, processing, and identification of emotional faces.
This repository provides the materials, datasets, and documentation for the Argentinian adaptation and validation of the Affective Norms for English Texts (ANET). The project includes affective ratings (Valence, Arousal, Dominance) and linguistic measures for a set of short emotional texts written in Rioplatense Spanish.
How does the relation between two words create humor? In this article, we investigated the effect of global and local contrast on the humor of word pairs. We capitalized on the existence of psycholinguistic lexical norms by examining violations of expectations set up by typical patterns of English usage (global contrast) and within the local context of the words within the word pairs (local contrast). Global contrast was operationalized as lexical-semantic norms for single-words and local contrast was operationalized as the orthographic, phonological, and semantic distance between the two words in the pair. Through crowd-sourced (Study 1) and best-worst (Study 2) ratings of the humor of a large set of word pairs (i.e., compounds), we find evidence of both global and local contrast on compound-word humor. Specifically, we find that humor arises when there is a violation of expectations at the local level, between the individual words that make up the word pair, even after accounting for violations at the global level relative to the entire language. Semantic variables (arousal, dominance, and concreteness) were stronger predictors of word pair humor whereas form-related variables (number of letters, phonemes, and letter frequency) were stronger predictors of single-word humor. Moreover, we also find that semantic dissimilarity increases humor, by defusing the impact of low-valence words-making them seem more amusing-and by enhancing the incongruence of highly imageable pairs of concrete words. (PsycInfo Database Record (c) 2022 APA, all rights reserved).
The research deals with the set of Serbian homonymous nouns (nouns with multiple unrelated meanings) presented in the norming study and in the visual lexical decision task experiment. Native speakers listed the meanings of homonymous words and provided word familiarity and word concreteness ratings. Accordingly, the first database of Serbian homonyms was constructed containing subjective meanings of homonymous nouns along with the estimated meaning probabilities, as well as a number of meanings, redundancy and entropy of the distribution of meaning probabilities, word familiarity and word concreteness. The processing disadvantage of homonymous nouns over unambiguous nouns was replicated in the visual lexical decision task. Additionally, the processing of homonymous nouns was linked with redundancy: the information theory measure of the balance of meaning probabilities. The results revealed that homonyms with higher redundancy of the meaning probability distribution (i.e., unbalanced meaning probabilities) were processed faster. This finding was in accordance with the hypothesis derived from the Semantic Settling Dynamics account of the processing of ambiguous words, according to which the competition among the unrelated meanings derived the processing disadvantage in homonymy. However, the same pattern was not observed for the number of meanings and entropy, inviting for further research of the processing of ambiguous words.
ABSTRACT An individual’s sense of the extent to which her or his body physically interacts with objects in the environment (body–object interaction; BOI) has been empirically shown to modulate lexical and semantic processing of object names. To allow for further exploration of the nature of those effects, BOI ratings for 750 Spanish nouns were obtained from 178 young adult participants. Statistical analyses showed moderate correlations between BOI indicators and some psycholinguistic indexes, such as word imageability and age of acquisition. In addition, an exploration of lexical associative relationships revealed that high-BOI words have a consistent tendency to be associated with words naming parts of the body. The ratings could be useful to researchers who are interested in manipulating or controlling for the effects of BOI in their language-processing studies. The complete norms are available for free downloading at Open Science Framework ( https://osf.io/kd5vf/ ).
Résumé Le but de la présente recherche est de contribuer à identifier les principes de l'organisation des représentations en mémoire. Nous avons collecté les productions d'exemplaires appartenant à 22 catégories sémantiques, dénotant soit des objets (naturels ou fabriqués), soit des activités humaines et ce à partir de deux consignes différentes: l'une insiste sur la production de mots, l'autre sur la production d'une image préalablement à la dénomination. Les résultats montrent: — une importante stabilité interindividuelle, permettant d'assigner un statut de « normes » à ces productions verbales; — une importante diversité entre catégories sur les critères que nous avons analysés; — la non-exclusivité des déterminations linguistiques (lexicales) dans la distribution des réponses des sujets; — la contribution possible des représentations imagées dans l'organisation de certaines catégories. Ces facteurs n'épuisant cependant pas la richesse des déterminants de la structuration catégorielle, ce travail suggère en conclusion de nouvelles investigations des déterminations de l'organisation des représentations cognitives humaines. Mots clefs: Représentation, catégorisation, lexique, imagerie.
Research on language and cognition relies extensively on psycholinguistic datasets or "norms". These datasets contain judgments of lexical properties like concreteness and age of acquisition, and can be used to norm experimental stimuli, discover empirical relationships in the lexicon, and stress-test computational models. However, collecting human judgments at scale is both time-consuming and expensive. This issue of scale is compounded for multi-dimensional norms and those incorporating context. The current work asks whether large language models (LLMs) can be leveraged to augment the creation of large, psycholinguistic datasets in English. I use GPT-4 to collect multiple kinds of semantic judgments (e.g., word similarity, contextualized sensorimotor associations, iconicity) for English words and compare these judgments against the human "gold standard". For each dataset, I find that GPT-4's judgments are positively correlated with human judgments, in some cases rivaling or even exceeding the average inter-annotator agreement displayed by humans. I then identify several ways in which LLM-generated norms differ from human-generated norms systematically. I also perform several "substitution analyses", which demonstrate that replacing human-generated norms with LLM-generated norms in a statistical model does not change the sign of parameter estimates (though in select cases, there are significant changes to their magnitude). I conclude by discussing the considerations and limitations associated with LLM-generated norms in general, including concerns of data contamination, the choice of LLM, external validity, construct validity, and data quality. Additionally, all of GPT-4's judgments (over 30,000 in total) are made available online for further analysis.
All words have properties linked to form, meaning and usage patterns which influence how easily they are accessed from the mental lexicon in language production, perception and comprehension. Examples of such properties are imageability, phonological and morphological complexity, word class, argument structure, frequency of use and age of acquisition. Due to linguistic and cultural variation the properties and the values associated with them differ across languages. Hence, for research as well as clinical purposes, language specific information on lexical properties is needed. To meet this need, an electronically searchable lexical database with more than 1600 Norwegian words coded for more than 12 different properties has been established. This article presents the content and structure of the database as well as the search options available in the interface. Finally, it briefly describes some of the ways in which the database can be used in research, clinical practice and teaching.
The primary goal of this project was to collect normative emotional valence and arousal ratings using the RADIATE facial database. The RADIATE database is one of the few that is racially diverse, yet it is underutilized, due in part to a lack of normative valence and arousal ratings. A secondary goal was to explore whether the race of the rater moderated emotion ratings. As part of an ongoing study, 204 participants (Asian: 9, Black: 25, Latinx: 39, White: 131) were randomly assigned to one of 10 blocks of 36 faces. Each block included faces counterbalanced on race, gender, and emotion so that each participant rated an identical number of faces with respect to these categories. Participants viewed faces in Qualtrics and rated each on valence (from 1-9, unpleasant to pleasant) and arousal (from 1-9, low to high). A 4-way Race of Rater x Race of Face x Emotion x Gender repeated-measures ANOVA with repeated-measures on the last 3 factors was used for valence and arousal ratings. As expected, across racial face categories, happy faces were rated as more pleasant (M = 6.50) and sad faces as more unpleasant (M = 3.03). In addition, happy (M = 4.29) faces were rated more emotionally arousing than sad (M = 3.76) and neutral faces (M = 3.29). The race of the rater moderated valence but not arousal ratings. Black raters rated Asian females as happier than Asian males and Latinx raters rated Latinas as sadder than Latinos, with no other evident effects. Present results contribute to sparse valence and arousal data for the RADIATE dataset. Results further suggest that emotional faces are not rated in a universal manner as some emotion theories presume. Implications of the results and future research directions are discussed.