1358 norm sets
Contains all data from this psycholinguistic study, including original surveys, compiled data, and analysis R code. Part of the toolkit of language researchers is formed of stimuli that have been rated on various dimensions. The current study presents modality exclusivity norms for 336 properties and 411 concepts in Dutch. Forty-two respondents rated the auditory, haptic, and visual strength of these words. Mean scores were then computed, yielding acceptable reliability values. Measures of modality exclusivity and perceptual strength were also computed. Furthermore, the data includes psycholinguistic variables from other corpora, covering length (e.g., number of phonemes), frequency (e.g., contextual diversity), and distinctiveness (e.g., number of orthographic neighbours), along with concreteness and age of acquisition. To test these norms, Lynott and Connell’s (2009, 2013) analyses were replicated. First, unimodal, bimodal, and tri-modal words were found. Vision was the most prevalent modality. Vision and touch were relatively related, leaving a more independent auditory modality. Properties were more strongly perceptual than concepts. Last, sound symbolism was investigated using regression, which revealed that auditory strength predicted lexical properties of the words better than the other modalities did, or else with a different direction. All the data and analysis code, including a web application, are available from https://osf.io/brkjw/. Data and analyses dashboard: https://pablobernabeu.shinyapps.io/Dutch-modality-exclusivity-norms/ (in case of downtime, please visit https://pablobernabeu.github.io/dashboards/Dutch-modality-exclusivity-norms/d.html). Online RStudio environment with data and code: https://mybinder.org/v2/gh/pablobernabeu/Modality-exclusivity-norms-747-Dutch-English-replication/master?urlpath=rstudio Paper (Bernabeu, 2018): https://psyarxiv.com/s2c5h The norms were used, and validated, in an experiment that implemented the conceptual modality switch. Data for that experiment may be found as a linked component of the present Project.
Research on metaphor has steadily increased over the last decades, as this phenomenon opens a window into a range of processes in language and cognition, from pragmatic inference to abstraction and embodied simulation. At the same time, the demand for rigorously constructed and extensively normed experimental materials increased as well. Here, we present the Figurative Archive, an open database of 997 metaphors in Italian enriched with rating and corpus-based measures (from familiarity to lexical frequency), derived by collecting stimuli used across 11 studies. It includes both everyday and literary metaphors, varying in structure and semantic domains. Dataset validation comprised correlations between familiarity and other measures. The Figurative Archive has several aspects of novelty: it is increased in size compared to previous resources; it includes a novel measure of inclusiveness, to comply with current recommendations for non-discriminatory language use; it is displayed in a web-based interface, with features for a flexible and customized consultation. We provide guidelines for using the Archive in future metaphor studies, in the spirit of open science.
In psycholinguistic research, careful selection and control of stimuli are essential for gaining insights into cognitive processes. Within this field, pictures often serve as stimuli, which requires the use of image databases to investigate linguistic, mnemonic, and visual perceptual phenomena in different populations (children without disabilities, adults, elderly people, illiterate, and brain-damage patients; see Soares et al., 2018 for more detail).Although several image databases provide norms for variables such as naming agreement (the most common name assigned to a picture by individuals; Snodgrass & Vanderwart, 1980), conceptual familiarity (the frequency with which individuals encounter or think about the depicted object; Snodgrass & Vanderwart, 1980) and visual complexity (judgments regarding the number of lines, intricacies, and details in an image; Snodgrass & Vanderwart, 1980; see also Székely & Bates, 2000 for an objective measure of picture visual complexity), these norms are available in multiple languages but are often restricted to a limited set of black-and-white line drawings (less than 300). Notably, such black-and-white images have been found to elicit weaker recognition compared to colored pictures (Sanfeliu & Fernandez, 1996;Rossion & Pourtois, 2004). In recent decades, there has been an increased effort to develop colored image datasets in various languages. However, many of them consider small datasets (usually less than 500 pictures, but see Brodeur et al., 2014;and Krautz & Keuleers, 2022, for more extensive datasets) and/or use different normalization protocols that complicate the process of comparing data and planning and executing cross-linguistic experiments (Soares et al., 2018;Duñabeitia et al., 2022;and Zhong et al, 2024 for overviews).The Multilingual Picture (MultiPic) database (Duñabeitia et al., 2022) was designed to address the limitations of previous databases by providing researchers with norms for naming agreement and concept familiarity for a set of colored images (500), selected from an initial pool of 750 images (Duñabeitia et al., 2018). To date, this database spans thirty-three languages, including American English, Australian English, Basque, Belgium Dutch, British English, Cantonese, Catalan, Cypriot Greek, Czech, Finnish, French, German, Greek, Hebrew, Hungarian, Italian, Korean, Lebanese Arabic, Malay, Malaysian English, Mandarin Chinese, Netherlands Dutch, Norwegian, Polish, European Portuguese, Quebec French, Rioplatense Spanish, Russian, Serbian, Slovak, Spanish, Turkish, and Welsh. The images depict specific concepts and general knowledge items, and the same data collection and preprocessing protocols were consistently applied across all languages.Expanding MultiPic to include additional languages, dialects, and language varieties worldwide would enable researchers to investigate lesser-studied languages beyond the predominant focus on English, facilitate direct cross-linguistic comparisons, and deepen our understanding of cognitive processes that are universal versus those that are language-specific. The primary objective of the present study was to norm MultiPic in Galician, a relatively under-researched language, enabling researchers to conduct studies with it. Galician is a Western Ibero-Romance language predominantly spoken in Galicia, an autonomous community in northwestern Spain, where it holds co-official status with Spanish.Psycholinguistic studies on Galician are less prevalent than those on Spanish, Portuguese, Catalan, and Basque (see Comesaña & Sá-Leite, 2024). This paucity of research investigating the specific cognitive mechanisms involved in both the comprehension and production of Galician is likely attributable to several factors, including the language's recent standardization in the 1980s (Pintos, 2025) and the scarcity of databases that would allow for the careful selection of linguistic materials for various experiments. The development of these tools would reinvigorate research on Galician by ensuring that experimental outcomes accurately mirror core cognitive processes. Consequently, they would provide essential scientific evidence to inform public policies related to Galician. This language, which coexists with Spanish, presents a distinctive opportunity to examine psycholinguistic theories of language processing within bilingual contexts.Although we generally adhere to the same data collection and preprocessing procedures described in the MultiPic database (Duñabeitia et al., 2022), several adaptations were necessary to accurately reflect Galician's linguistic reality. These modifications accounted for the linguistic diversity across the region, shaped by the contact between Galician and Spanish. For instance, because Galician remains subordinate to Spanish in many social contexts, speakers often incorporate Spanish words or adapted terms, with variations across different regions (cf. Rei-Doval, 2025). Thus, we considered the diverse linguistic varieties and regional differences within Galician, ensuring that the dataset represents the full spectrum of language use across different areas. This approach not only respects the sociolinguistic context of Galician but also allows for a more comprehensive understanding of the cognitive processes involved in bilingual language processing. That is, it will enable researchers to examine which cognitive processes are general and which are specific to different sociolinguistic contexts. Nevertheless, the experimental method, the preprocessing protocol, and the data structure are comprehensively detailed to provide researchers with the necessary framework for adapting the MultiPic to other languages with comparable sociolinguistic contexts.In conclusion, MultiPic and the Galician MultiPic, in particular, serve as valuable tools that enable researchers to design studies in Galician and other languages, where the properties of the materials have been rigorously tested in parallel.The complete dataset, including the data file, is publicly available in the following repositories: https://figshare.com/articles/dataset/Untitled_Item/19328939 and https://osf.io/ank4g/?view_only=43367d3dd27543b0aa66dfb8e71ce1fc.2 MethodParticipants were recruited over two months through social media and local newspaper advertisements. Their participation was voluntary. A total of 88 Galician speakers were initially recruited, surpassing the median sample size in the original MultiPic project (i.e., 80; Duñabeitia et al., 2022)). Still, three were excluded for not following task instructions (e.g., responding in a language other than Galician or basing answers on familiarity with the picture instead of its name).From the remaining 85 participants (47 women, 34 men, four preferred not to disclose their sex; mean age = 42 [age range of 18 to 82], SD = 19.05), around 28% of participants were from O Grove, 8% from Santiago de Compostela, 8% from A Coruña, 6% from Vigo, and the rest from 24 different places. Even though all were speakers of Galician, in their daily lives, around 28% spoke only Galician, 32% spoke more Galician than Castilian Spanish, 18% spoke both languages equally, 16% spoke more Castilian Spanish than Galician, and 6% spoke only Castilian Spanish. More than half had a university degree.We used the 500 colored pictures from the MultiPic database representing common concrete concepts. These pictures were in PNG format with a 300 × 300 pixels resolution and were initially created by a local artist commissioned by the authors of the original study (Duñabeitia et al., 2018). The set of 500 elements depicted was the same as those used in Duñabeitia et al. (2022), consisting of a pictorial set of digital line drawings derived from a list of imageable and concrete Spanish words taken from ESPAL (Duchon et al., 2013).The Galician MultiPic norming followed the standardized protocol of the original MultiPic project. Instructions were provided in Galician. Sociolinguistic data, including age, gender, number of languages spoken fluently, and possession of a university degree, were collected. However, unlike other languages, the sociolinguistic data for Galician was gathered in greater detail to ensure an accurate understanding of the sociolinguistic reality of the Galician language. Thus, questions were added regarding place of birth and place of residence, age of acquisition of Galician and Spanish, language balance, educational level, and socioeconomic status.Participants received a link and completed the tasks on their computers, tablets, or phones. However, fifteen elderly adults who were not computer literate gave their responses orally, which a team member transcribed. All participants completed the two tasks in the same order using the Gorilla Experiment Builder. First, participants were provided with a link and completed the tasks by typing their responses using a computer, tablet, or smartphone. However, fifteen elderly adults who were not computer literate provided their responses orally, which a team member transcribed. Considering this, all participants named each of the 500 randomly presented images, using no more than one word per concept. Then, they rated their familiarity with each concept on a 100-point scale, ranging from 0 (not familiar at all) to 100 (very familiar). If they did not know the name of an image, they could select the "?" button, which was recorded as an "I don't know" response. Before starting, participants completed two practice trials to familiarize themselves with the procedure. The experiment lasted approximately one hour, with breaks every 50 trials. Responses were coded to account for linguistic variations, including standard (i.e., the form accepted by the Real Academia Galega [Royal Galician Academy]) versus colloquial forms, dialectal differences, and influences from Spanish. To this end, and following preceding studies (Duñabeitia et al., 2018;2022), a native speaker of Galician reviewed and corrected spelling errors while also standardizing responses by merging basic variants of the same names (e.g., hyphenated or pluralized forms).A version of the Galician MultiPic is also available at https://figshare.com/articles/dataset/Untitled_Item/19328939. However, only the data considering the Galician nouns are provided here, even when the Galician noun was not the modal name (which occurred with 52 nouns [highlighted in yellow in the dataset provided in OSF]). Thus, the Galician MultiPic at Figshare includes nine columns (from A to I) as occurs with the other 33 languages of MultiPic, which corresponds to the Language provided, the Code (number of the picture), the Number of Responses, the H statistic, the Modal Response, the Modal Response Percentage, the "I don't know" Response Percentage, the Idiosyncratic Response Percentage, and the Familiarity.Sheet H-STATISTIC contains the calculation of the H-STATISTIC for each picture.Sheet CODEBOOK contains a detailed description of the information collected in both the raw and cleaned data frames used for analyses.Regarding the naming task, two measures were considered as in earlier Multipic studies: the mean H statistic and the mean modal response percentagefoot_0. These were analyzed, and the familiarity measures were recorded as well.The most notable finding is that the data exhibits an averaged H statistic of 0.71 and a mean modal response percentage of 73.56%. As mentioned above, only 52 pictures out of 500 had one unique response. The H statistic for Galician is higher than the average for MultiPic across 33 other languages (0.55). Indeed, only 5 out of 33 languages (Malay, Lebanese, Korean, Mandarin, and Cantonese) have higher H statistic values than Galician, and only 2 (Mandarin and Cantonese) lower mean modal response percentages (73.28% and 59.17%, respectively). Interestingly, of the official languages in the Iberian Peninsula, including Basque, Catalan, Galician, Portuguese, and Castilian Spanish, Galician exhibits the highest H statistic and the lowest mean modal response percentage. In comparison, the H statistic and the mean modal response percentage for Basque are 0.66 and 82.94%, for Catalan 0.45 and 88.98%, for Portuguese 0.37 and 90.38%, and for Castilian Spanish 0.30 and 93%, respectively. Table 1 summarizes the norms for the 500 images of the MultiPic in each of the languages of the Iberian Peninsula. The relatively high mean H statistic and low mean modal response percentages obtained in the current dataset suggest a higher lexical variability when compared to most languages included in the Multipic database and to the languages that coexist in the Iberian Peninsula.Correlation analyses on the H statistic and Familiarity values across languages of the Iberian Peninsula were conducted to validate individual dataset quality. We focus on comparative analyses in these languages because speakers share not only historical and linguistic connections but also cultural and educational influences that shape familiarity judgments. This is particularly relevant for Romance languages like Castilian Spanish, European Portuguese, Catalan, and Galician, which have significant lexical and structural similarities, as well as for Basque, which, despite being non-Romance coexists in the same sociolinguistic environment. While cross-linguistic correlations can occur even between typologically distant languages, as shown in previous studies, our focus here is on a more controlled linguistic and cultural space, allowing for a more precise interpretation of familiarity effects.A correlation analysis performed on the H statistic showed that all the Pearson pairwise correlation coefficients were significant at the p < 0.001 level, with r values ranging between 0.27 (Galician vs. European Portuguese) and 0.59 (Spanish vs. Catalan). The reason why r values between Galician and European Portuguese are the lowest despite their status as closely related languages with a shared medieval history as part of Galician-Portuguese, may be attributed to the distinct sociolinguistic contexts in which they have developed. These differing contexts have played a significant role in shaping lexical variation between the two languages. Note that Galician coexists with Castilian Spanish, a language of high prestige, which has led to the incorporation of numerous lexical borrowings from this language into Galician (see Dubert, 2025). Furthermore, the establishment of an official written standard for Galician did not occur until 1980, highlighting the relatively recent process of linguistic standardization. In contrast, Portuguese is the main language in Portugal and does not coexist with another widely spoken language, except for Mirandese, which is used in the specific region of Miranda do Douro. Additionally, Portuguese has a long-established linguistic tradition, with its first grammar and dictionary dating back to the 16th century. Since these early efforts, a strong normative tradition has been maintained (Santos, 2018).Likewise, the correlation analyses performed on different familiarity scores obtained for each item in each language showed that all the Pearson pairwise correlation coefficients were significant at the p < 0.001 level, with r values ranging between 0.79 (Catalan vs. Basque) and 0.83 (Galician vs. European Portuguese).Besides, a correlation analysis was conducted between the H statistic and Familiarity values in each official or co-official language from the Iberian Peninsula already tested in the MultiPic database. We found low to moderate negative and significant correlations in all of them. That is, the higher the values in familiarity, the lower the values in the H statistic, which makes sense as the higher the H statistic, the lower the name agreement. To be more precise, all Pearson pairwise correlation coefficients were significant at the p < 0.001 level, except for the Galician language ( p =.02), with r values ranging between -.10 (for the Galician) and -.44 (for the Catalan). The smallest correlation was found for Galician. At first, we thought that this was probably because it has greater lexical variability than the other languages. Indeed, if we look at the second language from the Iberian Peninsula that has a high lexical variability (Basque), we can see that it also showed a small correlation value between H statistic and Familiarity (-.25). However, when compared with other languages like Chinese or Malay that also have a great lexical variability we found high significant correlations (-.49 and -.90, respectively). Therefore, a more plausible explanation may lay on the fact that familiarity modulates agreement (and not the other way around). That is, if someone is not familiar (or that much familiar) with an object, they would be hesitant when naming it, and as a consequence, this would lead to lower agreement scores across participants. This would be true for all the languages. However, for Galician more variables than familiarity may be explaining this result such as the already mentioned coexistence with the Castilian language, the recent official written standard for Galician, which means that it is not perfectly implemented, and the desire of some people to reflect their dialectal variant. We recognize, however, that this is a tentative explanation that deserves further examination.Each variable's inter-rater reliability was determined by calculating intraclass correlations (ICCs) via a two-way random consistency model. ICCs revealed acceptable reliability for H statistic (ICC = 0.78 [0.75, 0.81]), and an excellent reliability for familiarity (ICC = 0.92 [0.91, 0.93]).Although all correlations were significant, findings underscore how social and regional factors, such as dialectal variation, hyper-Galician forms, and Spanish influence, shape lexical variation in Galician. A closer look at the responses given to each picture shows that some pictures had multiple interpretations that seem to reflect the variety of realities of the population, i.e., participants used nouns with different meanings (e.g., magnet vs. horseshoe), for example, by using different co-hyponyms or elements of the same semantic field ( figo vs. cebola [fig vs. onion]), or, in some cases, by focusing on different elements or areas of the image (e.g., for the picture of a shoulder, participants used names like costas, pel, or marrón, i.e., back, skin, or brown in English). On the other hand, some pictures were named with synonyms, such as xornal and periódico, two different Galician words to name a newspaper. Importantly, in many cases, the noun participants used depended on the dialectal variety of their region (e.g., vespa, avespa, and avéspora for wasp). Also, in some cases, participants used "hyper-Galician" forms, i.e., linguistic that are created when speakers to use they as or Galician. often these to the of Castilian Spanish, a shaped by the unique context of language contact and the relatively standardization of Galician in Galicia, as For example, the Galician form for is, but the hyper-Galician form is many of lexical from Spanish can be such as the Spanish word for or for In in more than of the images the modal form is the Spanish word (e.g., or or the adapted Spanish word (e.g., and in and Spanish, or and in and Spanish, may however, that the lexical variability in Galician is by the pictures than by the of the language as an In other the same concepts consistently the lowest agreement across languages. This does not seem to be the when calculating the mean H statistic of the Multipic database including all the languages tested this was = and the mean modal response percentage was = These values are to those provided in earlier studies with different of stimuli (e.g., et al., et al., et al., et al., et al., et al., and as Duñabeitia et al. have already out when comparing the data of languages, relatively low mean H statistic and the high mean modal response percentages of the current dataset suggest high name agreement across items, languages, and the materials for their use in different of experiments and these The of data from specific regions and the written of the experiment questions about and to the relatively recent standardization of the Galician language, many speakers are with the spelling or This questions about the of this task in a written Nevertheless, this the of comprehensive planning and cultural when the Galician MultiPic valuable insights for cross-linguistic studies and research, recognition of linguistic diversity while providing a framework for adaptations in other Galician of MultiPic a in psycholinguistic research by providing standardized norms for an language. This the lexical variation in Galician, shaped by regional and language contact with like regional and Multipic is a written than spoken the Galician MultiPic is an essential for cross-linguistic studies and the of bilingual cognitive valuable insights for research on language
The age of acquisition (AoA) refers to the age at which an individual learns specific items or words. Research on word recognition has shown that items with lower AoA—those acquired earlier—can be processed more quickly and accurately. In the field of Japanese word recognition, a large-scale database that would enable mega-study approaches has not been well-established. In this study, we developed an AoA norm for over 5,000 Japanese words. A total of 1,345 adults rated the AoA of 5,736 words using a 7-point scale. These ratings demonstrated satisfactory reliability. Furthermore, when examining the correlation with lexical decision task performance, words with lower AoA showed shorter response times and higher accuracy than words with higher AoA. These findings indicate that the AoA rating database collected in this study serves as a valuable resource for research using Japanese words.
Word Association Norms (WAN) are collections that present stimuli words and the set of their associated responses. The corpus is widely used in diverse areas of expertise. In order to reduce the effort to have a good quality resource that can be reproduced in many languages with minimum sources, a methodology to build Automatic Word Association Norms is proposed (AWAN). The methodology has an input of two simple elements: a) dictionary, and b) pre-processed Word Embeddings. This new kind of WAN is evaluated in two ways: i) learning word embeddings based on the node2vec algorithm and comparing them with human annotated benchmarks, and ii) performing a lexical search for a reverse dictionary. Both evaluations are done in a weighted graph with the AWAN lexical elements. The results showed that the methodology produces good quality AWANs.
Several lexical databases have been developed in both English-speaking countries and other countries, leading to numerous studies using these resources. A prominent example is the English Lexicon Project (ELP; Balota et al., 2007), a large-scale database containing behavioral data on English word processing. The ELP provides data for two main tasks: the lexical decision task (LDT) and the speeded naming task. Among these, the LDT is the most commonly utilized in word-processing research, largely because 1) it is easy to implement, and 2) it can be conducted online with relative ease (Lieber et al., 2014).In the LDT, participants are asked to decide as quickly and accurately as possible whether a visually presented string of letters forms a real word or a non-word. By analyzing the response time from when the string is presented until the participant makes a decision, researchers can evaluate the speed of word access and semantic processing. The LDT has been employed not only to assess word processing efficiency and cognitive load but also to investigate the structure of the mental lexicon and concept representation. For example, researchers have examined the relationship between LDT response times and various word properties, including the frequency effect, where more frequent words are processed faster and reexamined using the LDT data (Brysbaert et al., 2011).Numerous psycholinguistic studies have explored the semantic properties of word recognition using LDT data. Recently, LDT databases have expanded beyond English, with resources available in languages such as Chinese (Tse et al., 2017), French (Ferrand et al., 2017), and Spanish (Aguasvivas et al., 2018), allowing for more efficient research across languages. For example, researchers have tested hypotheses involving grounded cognition and embodied cognition (Barsalou, 2008) in word recognition and explored the relationship between word recognition and sensorimotor information across various languages, e.g., English (Pexman et al., 2019; Sidhu et al., 2014), French (Lalancette et al., 2024), and Spanish (Alonso et al., 2018). They further examined theoretical predictions with large-scale survey data, often using lexical decision task (LDT) reaction times as the dependent variable in regression analyses. While earlier findings have supported these theories by showing consistent trends across languages, recent discussions have highlighted cross-linguistic variability in these effects (Alonso et al., 2018; Lalancette et al., 2024). Such hypothesis testing using a database reduces stimulus bias by incorporating many words (see Dymarska et al., 2023) and enables new discoveries through cross-linguistic comparisons.In Japanese, several databases are available, as will be discussed later. For example, databases exist for attributes such as word imageability (Sakuma et al., 2005) and familiarity (Asahara, 2020), each containing evaluative data for tens of thousands of words. These databases have long been used in various ways, such as serving as control variables in numerous Japanese word recognition studies (e.g., Mizuno and Matsui, 2018; Mochizuki and Ota, 2020, 2024). However, no LDT database currently exists for Japanese, posing a challenge to psycholinguistic research on the Japanese language as a result of limited resources. Of course, lexical decision tasks have been widely used in Japanese word recognition studies (e.g., Kawakami, 2002; Kusunose et al., 2013). However, the number of stimulus words used in these studies is significantly smaller compared to databases such as the ELP (Balota et al., 2007). Furthermore, the data are not always publicly available, which limits their utility as resources. Given the increasing emphasis on cross-linguistic validation—particularly in studies of abstract concepts shaped by language and culture (Dove, 2018)—developing a large-scale Japanese LDT database would not only aid Japanese researchers but also contribute to the broader field. Therefore, this study aimed to construct a Japanese version of LDT database.It is important to note that individual differences in LDT response times exist (e.g., Hawker and Ferraro, 2007; Yates and Slattery, 2019; Lim et al., 2020). To enhance the database, we collected data on participants’ individual characteristics following the LDT. Specifically, participants completed the ENDCOREs, which measures interpersonal communication skills (Fujimoto and Daibo, 2007), and the Japanese version of the Plymouth Sensory Imagery Questionnaire (Psi-Q) (Fukui and Aoki, 2022). The ENDCOREs assesses six dimensions of communication: self-control, expressiveness, comprehension, assertiveness, acceptance of others, and relational adjustment. The Psi-Q evaluates the vividness of mental imagery across sensory modalities (i.e., vision, sound, smell, taste, touch, body, and emotion), capturing individual differences in multisensory imagery.Although we do not hypothesize a direct relationship between these individual difference variables and simple LDT response times (e.g., the higher/lower a score, the slower/faster the response time), they may serve as possible predictors for validating certain content. For instance, the grounded or embodied cognition framework (Barsalou, 2020, 2008) posits that processing words or concepts involves simulating the sensory modalities through which they are acquired. Consistent with this, processing words rich in sensorimotor information tends to be more efficient (Lynott et al., 2020; Siakaluk et al., 2008; Sidhu et al., 2014; Sidhu and Pexman, 2016; Tillotson et al., 2008). Individual differences in sensitivity to sensory and motor modalities may interact with word characteristics and influence LDT performance. Furthermore, the “Words as Social Tools” (WAT) perspective (Borghi and Binkofski, 2014) posits that simulating social and linguistic information is crucial for understanding abstract concepts (Borghi et al., 2019). Therefore, words with a stronger social nature may be processed more efficiently (Diveica et al., 2023), and the interaction between verbal sociality and individual sociality may affect LDT response times. Since ENDCOREs reflect an individual’s communication skills, individuals with high social interaction skills may find it easier to simulate socially relevant words. Consequently, they might be more efficient in processing abstract words with strong social characteristics. While the present study did not specifically examine the relationship between individual differences and LDT response times, future research could benefit from incorporating these variables into the database.This report introduces the Japanese LDT database (JALEX), which incorporates individual differences among participants. The response time and accuracy data can be used for future psycholinguistic studies involving Japanese participants. Additionally, while no hypotheses were tested, future research may explore the role of individual differences as needed.In the development of psycholinguistic norms, approximately 30 to 40 observations per word are typically required (Balota et al., 2007; Ferrand et al., 2017). However, we recruited a relatively large number of participants to account for potential dropouts, as this was an online study, and to develop more reliable norms.Participants were recruited through a crowdsourcing service Yahoo! Crowdsourcing (https://crowdsourcing.yahoo.co.jp/). A total of 2,689 individuals accessed the task. However, 1,037 either did not start, failed to complete the task, or provided no responses. Ultimately, 1,652 participants completed the task. All participants self-reported as native Japanese speakers. Among them, 1,226 were men, 407 were women, two identified as other genders, and 17 chose not to respond. The mean age was 51.07 years (SD = 11.91), with a range from 18 to 85 years. The participants’ highest levels of education were as follows: 26 had completed doctoral programs, 119 had master’s degrees, 1,069 were college graduates, 21 had finished high school, 28 had completed junior high school, and 26 chose not to respond. As detailed below, the words were divided into 38 lists. With 1,652 participants, this resulted in approximately 43 participants per list. To develop JALEX databases, we selected words with semantic properties listed in multiple extant databases (DBs). This approach ensured consistency with previous word recognition studies and supported continuity in future research. We followed a specific selection procedure. First, we used the Word List by Semantic Principles, revised and enlarged edition (WLSP, National Institute for Japanese Language and Linguistics, 2004) as the master list. From this, we selected words that appeared in all eight of the following DBs: the word familiarity DB (Asahara, 2020), an alternate word familiarity DB (Fujita and Kobayashi, 2020), the word frequency DB (Amano and Kondo, 2000), the NINJAL-LWP for TWC word frequency DB (University of Tsukuba et al., 2013), the word difficulty DB (Kajiwara et al., 2020), the imageability DB for visual words (Sakuma et al., 2005), the semantic orientations DB (Takamura et al., 2005), and the abstractness DB for Japanese words (The Social Computing Laboratory, 2021). Following this procedure, we selected 5,736 Japanese words as stimuli. These included 4,977 nouns, 648 verbs, and 111 adjectives.For each word, linguistic characteristics such as orthographic neighborhood size (ONS), phonological neighborhood size (PNS), orthographic Levenshtein distance 20 (OLD20, Yarkoni et al., 2008), the number of letters, and the number of morae (a rhythmic unit of sound) were calculated. The PNS was computed by decomposing the ‘phonetic’ (読み) variable in the WLSP (National Institute for Japanese Language and Linguistics, 2004) by mora and calculating how many words in the WLSP had one mora replaced. Similarly, the ONS was calculated by decomposing the ‘letter (見出し本体)’ variable in the WLSP into individual characters and determining how many words had one letter replaced. OLD20 was calculated using the old20 function in the vwr package (Keuleers, 2013) in R (R Core Team, 2022), based on the ‘letter (見出し本体)’ variable in the WLSP.In addition, non-words were constructed as filler items for the LDT. First, from the WLSP, we excluded words with one mora, words containing spaces, symbols, or particles, homophones, and items with repetitive morae (e.g., ha-ha-ha [ha/ha/ha]), as these could not be transformed into non-words using the procedure described below. The remaining items were then decomposed into morae, and each mora was randomly shuffled. If the resulting item was not found in the WLSP, it was considered a non-word candidate. This process yielded 63,305 non-word candidates, from which we randomly selected 5,736 to serve as fillers for the LDT. The authors reviewed these candidates, and those deemed too similar to real words were replaced with different non-word candidates. All non-word stimuli are available for reference on Open Science Framework (OSF).The words and non-words were randomly divided into 38 lists, each containing 150 or 151 words (150 × 2 + 151 × 36 = 5,736) with an equal number of non-words.The LDT task was conducted online, and participants accessed the LDT program via their own PCs. The program was created using PsychoPy (Peirce et al., 2019) and hosted on Pavlovia (https://pavlovia.org/). After obtaining informed consent from the participants, they were instructed to begin the task. In the LDT, a blank screen appeared for 200 ms, followed by a fixation point in the center of the screen for 300 ms. A string of characters was then presented, and participants had to decide as quickly and accurately as possible whether the string represented a real Japanese word. The string remained on the screen until a response was made or for up to 2,000 ms. Participants pressed the ‘L’ key for words and the ‘S’ key for non-words. If the response was correct, the task proceeded to the next trial; if incorrect, a feedback message ("Wrong") appeared in red. If no response was given within 2,000 ms, the feedback message ("Too late") was displayed in red for 300 ms. Words and non-words were presented in random order. Participants completed 20 practice trials before starting the actual task. The practice trials used different stimuli from those in the actual task.During the task, participants were allowed to take a break for a maximum of 60 seconds between the 100th and 200th trials. During the break, their percentage of correct answers was displayed to encourage them to continue. Upon completing the LDT, participants answered the ENDCOREs (Fujimoto and Daibo, 2007) and Psi-Q (Fukui and Aoki, 2022) questionnaires. Additionally, they provided demographic information, including gender, age, dominant hand, highest level of education, and native language. Data were collected on May 20 and May 21, 2024. We calculated the accuracy rate for each participant, and the lowest percentage of correct responses exceeded 75%. Since no participants demonstrated an exceptionally low accuracy rate, data from all participants were retained for analysis.The procedure for processing the response time data followed that employed in ELP (Balota et al., 2007). First, we extracted only correct trials, where the "L" key was pressed for word stimuli, and excluded any trials with response times below 200 ms. Second, we removed trials that deviated by ±3 SD from the participant’s mean response time. This resulted in the exclusion of 1.96% of trials as outliers.The distribution of response times averaged by item is showed in Figure 1. We presented the partial correlations with existing DB variables referenced in stimulus selection to examine the convergent validity of the response time data (Figure 2). These findings confirmed the phenomena predicted in prior studies. Specifically, we confirmed the frequency effect (Rubenstein et al., 1971) and familiarity effect (Connine et al., 1990), where lexical decision times decrease as word frequency and familiarity increase. We also observed the imageability effect (Balota et al., 2004), where higher imageability leads to faster lexical decisions, and the orthographic similarity effect (Yarkoni et al., 2008), where greater Levenshtein distance results in longer response times. While few studies have reported simple or partial correlations with these variables in Japanese, several experimental studies using Japanese words as stimuli have observed effects similar to those identified in the present study. For instance, Japanese word recognition research has reported faster word processing for words with higher imageability (Ogawa and Nittono, 2018) and higher frequency (Mizuno and Matsui, 2015). The relationship between response time and Kajiwara’s (2020) difficulty rating has yet to be investigated. However, it is reasonable to predict that more difficult words would require longer processing times for comprehension. The current analysis identified a slight positive correlation between word difficulty and response time. These findings suggest that JALEX is valid to a considerable extent.The partial correlations between familiarity, semantic orientation, abstractness, and response time were significant, but the effects were small. Of these, the zero-order correlation for familiarity was r = -.42, suggesting that higher familiarity facilitates responses when not adjusted for covariates. Zero-order correlations for abstractness revealed a small effect (r =.11), indicating that processing was slightly suppressed for more abstract words, consistent with the representativeness effect (Cortese and Balota, 2012). This study also found that words with high ONS had shorter lexical decision times. The results showed that high ONS words took less time to judge than low ONS words when using Kanji words (Mizuno and Matsui, 2014), which is consistent with the current results. However, when using Katakana words, the inhibitory effect was observed, indicating that low ONS words took less time to judge than high ONS words (Kawakami, 2002). Furthermore, an interaction between ONS and PNS has also been observed in lexical decision performance for katakana words (Hino et al., 2011). The difference in these results may be caused by the limited number of words used in the experiment and the factors of the orthographic form. In studies using word norms, it is particularly important to consider the extent of word coverage and the absence of bias (Dymarska et al., 2023). The failure to replicate the effects observed in previous studies in the present analysis of a relatively large database may be attributed to biases in the stimulus sets used in those studies, which could have significantly influenced their results. Future research should assess the reproducibility of findings from previous studies by leveraging large databases, such as JALEX, and conducting comprehensive analyses.In the present study, the imageability effect (Balota et al., 2004) was replicated even after controlling for linguistic statistical variables, such as the frequency of neighboring words. The effects of psycholinguistic variables, such as the imageability effect, are often discussed in relation to semantic richness (Pexman et al., 2013). Semantic richness refers to the idea that words associated with more semantic information have richer semantic representations, enabling them to be processed more quickly and accurately. In other words, our study replicates in Japanese the finding that the ease of forming a mental image is an important semantic variable in word representations. As discussed in the introduction, the relationship between sensorimotor information and word recognition is explained by the concept of semantic richness—specifically, the richness of the semantic dimension of sensorimotor information facilitates word recognition. In future studies, it will be important to investigate the nature of concept representations by examining psycholinguistic variables influencing word recognition beyond imageability.Furthermore, this study is the first DB of LDT to include individual difference variables for respondents, paving the way for future research on individual differences using JALEX. In word recognition research, it has been observed that certain words exhibit significant individual differences and high variability in ratings of psychological variables (Paisios et al., 2023). A key limitation of the previous DB of LDT is that they did not provide individual difference data for participants, making them unsuitable for studying individual differences in words with high variability in ratings among individuals. Future research using JALEX is expected to refine further grounded cognition theory (Barsalou, 2020, 2008) and advance WAT theory (Borghi and Binkofski, 2014), particularly by promoting individual difference studies on the simulation of sensorimotor information and those related to social communication.This database represents the most comprehensive dataset on the efficiency of Japanese visual word processing and stands as a powerful resource for future research in psychology and linguistics. A unique feature of this database is its inclusion of individual difference variables for participants, allowing researchers to analyze these differences in future studies.However, it is important to note that some words in the dataset had lower accuracy rates. For example, at least 15 items had a correct response rate below 70%, with fewer than 20 observations. Items with fewer observations may exhibit lower reliability and reproducibility compared to others. While we did not exclude these items in the current analysis, researchers should be mindful of their presence when using the database.All data reported in this study can be found in the OSF Repository (https://osf.io/qr2sg). Information from existing databases used for validation cannot be included in the data resources of this study as a result of copyright, however, such information is available in the literature.
L’objectif premier de ce travail etait de caracteriser les images proposees par Bonin, Peereman, Malardier, Meot et Chalard (2003) en termes d’âge d’acquisition (AoA) objectif, recueilli aupres d’enfants âges de 2: 6 a 10: 11 ans. (http:// www. unice. fr/ LPEQ/ base_ AoA/ aoa_ intro. php). La comparaison avec les normes en francais des images de Snodgrass et Vanderwart (1980) montre que les nouvelles images correspondent, en moyenne, a des mots plus rares et d’AoA plus tardif. Cependant, correlations et regressions multiples indiquent que les memes facteurs sont impliques dans l’emergence de l’AoA (variabilite d’imagerie et frequence lexicale) a travers les deux jeux d’images. De plus, quelle que soit la base d’images, l’AoA est significativement mieux correle et specifiquement predit par les frequences lexicales du vocabulaire de l’enfant (NOVLEX, MANULEX) que par les frequences calculees sur des corpus du vocabulaire de l’adulte (BRULEX, LEXIQUE).
Extending psycholinguistic research into the lexical representation of two-kanji compound words within the Japanese mental lexicon (Joyce, 2002, 2004), this paper reports on a large-scale word association survey for basic Japanese vocabulary. The database of word association norms, which is being compiled from various survey formats including a web-based version of the survey, supplements existing databases concerning the lexical features of Japanese vocabulary (Amano & Kondo, 1999; Yokoy m a a, Sasahara, Nozaki & Long, 1998), such as familiarity ratings and frequency counts, which are essential for cognitive science research. A particularly promising application of the word association norms data, however, is the creation of lexical association network maps that capture important proprieties of words and their interconnectivity. These maps complement other approaches that attempt to tap into aspects of lexical knowledge, such as WordNet, thesauri, ontologies, and collocation data, while avoiding some of their problems. There are also direct and interesting lexicographical and Japanese language learning applications of the
Résumé Cet article présente des normes d’imageabilité (ou valeurs d’imagerie) pour un ensemble de 1493 mots. Des analyses statistiques réalisées sur ces normes révèlent une fidélité élevée. Les scores d’imageabilité se révèlent par ailleurs assez modestement corrélés avec d’autres variables psycholinguistiques (par ex., fréquences lexicales, âge d’acquisition). Des analyses restreintes à un sous-échantillon de mots en français, ainsi que d’autres sur des mots normés pour l’anglais, révèlent que le nombre de traits sémantiques est modérément positivement corrélé aux scores d’imageabilité, contrairement à l’hypothèse selon laquelle la richesse sémantique est adéquatement indexée par l’imageabilité.
Abstract The functional variants of International English are often differently distributed in the different regional standards. With evidence from the corpus of Australian English, this has already been shown for lexical variants such as will/shall, maybe/perhaps etc. In this paper evidence from the Australian corpus is used to discuss a number of variables in a) morphology b) the system of conjunction c) the system of quantifiers. The redistribution of morphological variants-edl-t (as in burned/burnt), and -wards(s) (as in downward(s)) showed a tendency to assign different grammatical roles to each variant. Among the conjunctions, apart from individual differences the most interesting finding was the higher level overall in the use of subordinating conjunctions, when Australian newspaper data was compared with the equivalent in Britain or America. A possible explanation for this invokes the Hallidayan principle that subordination is actually more common in speech than in writing. The suggestion is that Australian press reporting approximates more closely to spoken than to written norms of language. But on the quantifiers a few/several the corpus provides no support for a new popular use of several, to mean vaguely large number.
Our internal repository of words, often known as the mental lexicon, has primarily been modelled by psychologists as some kind of network. One way to probe its organisation and access mechanisms is by means of word association techniques, which have rarely been applied on Chinese. This paper reports on the design and implementation of a pilot word association test on native Hong Kong Cantonese speakers. The test contains 500 stimulus words, carefully selected and controlled on important factors including word frequency, part-of-speech, syllabicity, concreteness and vocabulary type. The resulting association norms based on 58 participants reveal interesting properties of the Chinese mental lexicon, such as the dominance of disyllabic and nominal concepts, and collocational associations. Despite its current small scale, the word association norms obtained from this study do not only offer first-hand psycholinguistic evidence for investigating the Chinese mental lexicon but also provide a useful resource to inform future studies in Chinese lexical access, lexical semantics and lexicography. 1
Investigation of affective and semantic dimensions of words is essential for studying word processing. In this study, we expanded Tse et al.'s (Behav Res Methods 49:1503-1519, 2017; Behav Res Methods 55:4382-4402, 2023) Chinese Lexicon Project by norming five word dimensions (valence, arousal, familiarity, concreteness, and imageability) for over 25,000 two-character Chinese words presented in traditional script. Through regression models that controlled for other variables, we examined the relationships among these dimensions. We included ambiguity, quantified by the standard deviation of the ratings of a given lexical variable across different raters, as separate variables (e.g., valence ambiguity) to explore their connections with other variables. The intensity-ambiguity relationships (i.e., between normed variables and their ambiguities, like valence with valence ambiguity) were also examined. In these analyses with a large pool of words and controlling for other lexical variables, we replicated the asymmetric U-shaped valence-arousal relationship, which was moderated by valence and arousal ambiguities. We also observed a curvilinear relationship between valence and familiarity and between valence and concreteness. Replicating Brainerd et al.'s (J Exp Psychol Gen 150:1476-1499, 2021; J Mem Lang 121:104286, 2021) quadratic intensity-ambiguity relationships, we found that the ambiguity of valence, arousal, concreteness, and imageability decreases as the value of these variables is extremely low or extremely high, although this was not generalized to familiarity. While concreteness and imageability were strongly correlated, they displayed different relationships with arousal, valence, familiarity, and valence ambiguity, suggesting their distinct conceptual nature. These findings further our understanding of the affective and semantic dimensions of two-character Chinese words. The normed values of all these variables can be accessed via https://osf.io/hwkv7.
The research deals with the set of Serbian homonymous nouns (nouns with multiple unrelated meanings) presented in the norming study and in the visual lexical decision task experiment. Native speakers listed the meanings of homonymous words and provided word familiarity and word concreteness ratings. Accordingly, the first database of Serbian homonyms was constructed containing subjective meanings of homonymous nouns along with the estimated meaning probabilities, as well as a number of meanings, redundancy and entropy of the distribution of meaning probabilities, word familiarity and word concreteness. The processing disadvantage of homonymous nouns over unambiguous nouns was replicated in the visual lexical decision task. Additionally, the processing of homonymous nouns was linked with redundancy: the information theory measure of the balance of meaning probabilities. The results revealed that homonyms with higher redundancy of the meaning probability distribution (i.e., unbalanced meaning probabilities) were processed faster. This finding was in accordance with the hypothesis derived from the Semantic Settling Dynamics account of the processing of ambiguous words, according to which the competition among the unrelated meanings derived the processing disadvantage in homonymy. However, the same pattern was not observed for the number of meanings and entropy, inviting for further research of the processing of ambiguous words.
How does the relation between two words create humor? In this article, we investigated the effect of global and local contrast on the humor of word pairs. We capitalized on the existence of psycholinguistic lexical norms by examining violations of expectations set up by typical patterns of English usage (global contrast) and within the local context of the words within the word pairs (local contrast). Global contrast was operationalized as lexical-semantic norms for single-words and local contrast was operationalized as the orthographic, phonological, and semantic distance between the two words in the pair. Through crowd-sourced (Study 1) and best-worst (Study 2) ratings of the humor of a large set of word pairs (i.e., compounds), we find evidence of both global and local contrast on compound-word humor. Specifically, we find that humor arises when there is a violation of expectations at the local level, between the individual words that make up the word pair, even after accounting for violations at the global level relative to the entire language. Semantic variables (arousal, dominance, and concreteness) were stronger predictors of word pair humor whereas form-related variables (number of letters, phonemes, and letter frequency) were stronger predictors of single-word humor. Moreover, we also find that semantic dissimilarity increases humor, by defusing the impact of low-valence words-making them seem more amusing-and by enhancing the incongruence of highly imageable pairs of concrete words. (PsycInfo Database Record (c) 2022 APA, all rights reserved).
ABSTRACT An individual’s sense of the extent to which her or his body physically interacts with objects in the environment (body–object interaction; BOI) has been empirically shown to modulate lexical and semantic processing of object names. To allow for further exploration of the nature of those effects, BOI ratings for 750 Spanish nouns were obtained from 178 young adult participants. Statistical analyses showed moderate correlations between BOI indicators and some psycholinguistic indexes, such as word imageability and age of acquisition. In addition, an exploration of lexical associative relationships revealed that high-BOI words have a consistent tendency to be associated with words naming parts of the body. The ratings could be useful to researchers who are interested in manipulating or controlling for the effects of BOI in their language-processing studies. The complete norms are available for free downloading at Open Science Framework ( https://osf.io/kd5vf/ ).
Résumé Le but de la présente recherche est de contribuer à identifier les principes de l'organisation des représentations en mémoire. Nous avons collecté les productions d'exemplaires appartenant à 22 catégories sémantiques, dénotant soit des objets (naturels ou fabriqués), soit des activités humaines et ce à partir de deux consignes différentes: l'une insiste sur la production de mots, l'autre sur la production d'une image préalablement à la dénomination. Les résultats montrent: — une importante stabilité interindividuelle, permettant d'assigner un statut de « normes » à ces productions verbales; — une importante diversité entre catégories sur les critères que nous avons analysés; — la non-exclusivité des déterminations linguistiques (lexicales) dans la distribution des réponses des sujets; — la contribution possible des représentations imagées dans l'organisation de certaines catégories. Ces facteurs n'épuisant cependant pas la richesse des déterminants de la structuration catégorielle, ce travail suggère en conclusion de nouvelles investigations des déterminations de l'organisation des représentations cognitives humaines. Mots clefs: Représentation, catégorisation, lexique, imagerie.
Résumé Les mots homonymes (par exemple, « avocat ») sont largement utilisés dans des expériences en psychologie cognitive afin d'étudier le traitement du langage, la levée des ambiguïtés lexicales et l'organisation en mémoire des représentations lexicales et sémantiques. Ces expériences requièrent le contrôle des relations associatives entre différents stimulus et des fréquences relatives des différentes acceptions des mots homonymes. L'objectif principal des normes d'associations verbales que nous présentons est de permettre aux chercheurs de réaliser de tels contrôles. Chacun des 162 items ambigus a été présenté à 100 sujets dans une tâche d'association libre. La totalité des réponses est présentée, ainsi que la fréquence relative des acceptions estimées à partir de ces normes. Mots clés: normes d'association libre, ambiguïté lexicale, homonymie, fréquence relative des acceptions.
Tangram pictures are abstract pictures which may be used as stimuli in various fields of experimental psychology and are often used in the field of dialogue psychology. The present study provides the first norms for a set of 332 tangram pictures. These pictures were standardized on a set of variables classically used in the literature on cognitive processes, such as visual perception, language, and memory: name agreement, image agreement, familiarity, visual complexity, image variability, and age of acquisition. Furthermore, norms for concreteness were also provided owing to the influence of this variable on the processes involved in lexical production. Correlational analyses on all variables were performed on the data collected from French native speakers. This new set of standardized pictures constitutes a reliable database for researchers when they select tangram pictures. Given the abstract nature of tangram pictures, this paper also discusses the similarities and differences with the literature on line drawings, and highlights their value for dialogue psychology studies, for psycholinguistics studies, and for cognitive psychology in general.
Norms of rated subjective frequency of use and imagery on seven-point scales are reported for 1,916 French nouns. Subjective frequency was defined as the rated frequency of occurrence of words in spoken French, and imagery was defined as the rated case with which a word aroused a mental image. The mean, standard deviation, and percentile rank of the frequency and imagery ratings for each item are presented in the Appendix together with their objective frequency of occurrence in Baudot's (1992) dictionary. Interjudge reliability was assessed by calculating the correlation between the mean ratings of items repeated in the booklet, between the mean ratings obtained from odd-numbered and even-numbered respondents, and by computing the Cronbach alpha statistic for each page of the booklet. These reliability estimates were equal to or greater than.92 for frequency and for imagery, confirming the high level of interjudge consistency. Although the estimates provided by female and male participants were highly correlated (r =.97), the former gave a slightly higher frequency rating to the word sample but a slightly lower imagery rating than the latter did. Moreover, female respondents gave slightly more extreme ratings on the frequency and imagery scales. An analysis of the absolute difference between female and male ratings revealed a discrepancy of one half point or more on 20% of the word sample for frequency and 13% for imagery. On both scales, the mean absolute difference between male and female ratings was larger than that obtained by chance alone. This finding highlights the possibility that some words may not be equally familiar to women and men or may not evoke imagery with the same ease in these groups. Validity estimates for the frequency and imagery ratings were derived from correlations with scale values drawn from other normative studies. These correlation coefficients were equal to or greater than.78 for frequency and.86 for imagery, confirming the high level of consistency between this and other studies. An analysis of the relationship between subjective frequency and imagery ratings indicated that these variables are generally uncorrelated but exceptions occur. In the present study the coefficient of the correlation between subjective frequency and imagery was.24. However, when items with extreme mean frequency were excluded from the calculation, the correlation coefficient dropped to.04 and was no longer significant. Imagery ratings from five independent studies were all positively and significantly correlated with Vikis-Freibergs's (1974) frequency estimates, which were obtained from a free-association task. This finding suggests that word association, as a form of cued recall, may be influenced by several stimulus attributes including prior frequency of association and imagery-evoking value. The pattern of correlation between imagery ratings and text-based frequency estimates is not coherent. It reveals significant correlations only in select cases and no consistent polarity of linear relationship. The main contribution of this research is to provide reliable estimates of subjective frequency and imagery value for a word sample that is larger than those included in previous studies. A close examination of the linear relationship among the various sources of frequency and imagery data underscores the risk of confounding these variables in the selection of lexical stimuli for research.
Concreteness is a fundamental dimension of word semantic representation that has attracted more and more interest to become one of the most studied variables in the psycholinguistic and cognitive neuroscience literature in the last decade. Concreteness effects have been found at both the brain and the behavioral levels, but they may vary depending on the constraints of the context and task demands. In this study, we collected concreteness norms for English and Italian words presented in different context sentences to allow better control and manipulation of concreteness in future psycholinguistic research. First, we observed high split-half correlations and Cronbach's alpha coefficients, suggesting that our ratings were highly reliable and can be used in Italian- and English-speaking populations. Second, our data indicate that the concreteness ratings are related to the lexical density and accessibility of the sentence in both English and Italian. We also found that the concreteness of words in isolation was highly correlated with that of words in context. Finally, we analyzed differences between nouns and verbs in concreteness ratings without significant effects. Our new concreteness norms of words in context are a valuable source of information for future research in both the English and Italian language. The complete database is available on the Open Science Framework (doi: 10.17605/OSF.IO/U3PC4).
In this paper, two word association (WA) studies are presented in support of recent arguments against the use of native-speaker (NS) norms in WA research. In Study 1, first-language (L1) and second-language (L2) WA norms lists were developed and compared to learner responses as a means of measuring L2 proficiency. The results showed that L2 norms provided a more sensitive measure of L2 lexical development than did traditional NS norms. Study 2 was designed to test the utility of native norms databases in predicting the primary WA responses of Japanese learners to high-frequency English cues. With the exception of only extremely frequent cues, it was shown that native norms were not successful in predicting learner responses. The results of both studies are discussed in terms of cultural and linguistic differences, geographic distance, and dissimilarities in word knowledge between respondent populations. Finally, a proposal is made for the construction of a Japanese WA database of English responses (J-WADE). The methods by which it will be developed, key features, and employment in future research are outlined.
Large-scale word association datasets are both important tools used in psycholinguistics and used as models that capture meaning when considered as semantic networks. Here, we present word association norms for Rioplatense Spanish, a variant spoken in Argentina and Uruguay. The norms were derived through a large-scale crowd-sourced continued word association task in which participants give three associations to a list of cue words. Covering over 13,000 words and +3.6 M responses, it is currently the most extensive dataset available for Spanish. We compare the obtained dataset with previous studies in Dutch and English to investigate the role of grammatical gender and studies that used Iberian Spanish to test generalizability to other Spanish variants. Finally, we evaluated the validity of our data in word processing (lexical decision reaction times) and semantic (similarity judgment) tasks. Our results demonstrate that network measures such as in-degree provide a good prediction of lexical decision response times. Analyzing semantic similarity judgments showed that results replicate and extend previous findings demonstrating that semantic similarity derived using spreading activation or spectral methods outperform word embeddings trained on text corpora.
Sensorimotor information is vital to the conceptual representation of our knowledge system. This study collects perceptual and action ratings for 664 disyllabic nouns among 438 native speakers and creates the first and largest dataset of sensorimotor norms for nouns in Chinese. Using aggregated semantic covariates, including concreteness ratings from a concreteness rating study, as well as the reaction times and error rates from a lexical decision study, our current work demonstrates the strengths of sensory modalities and action effectors in Chinese nouns and explores the contributions of embodied experiences in reflecting orthographic representations and semantic processing in the Chinese language. This study contributes valuable data sources to the study of Chinese lexical processing and highlights the importance of sensorimotor information and embodied manifestations in the semantic representations of concepts. Our results also support the language universal that orthographic awareness in lexical processing and reading supersedes phonological awareness.
Several norms of psycholinguistic features of Chinese characters exist in Mandarin Chinese, but only a few are available in Cantonese or in the traditional script, and none includes semantic radical transparency ratings. This study presents subjective ratings of age-of-acquisition (AoA), familiarity, imageability, concreteness, and semantic radical transparency in 4376 Chinese characters. The single Chinese characters were rated individually on the five dimensions by 20 native Cantonese speakers in Hong Kong to form the Hong Kong Chinese Character Psycholinguistic Norms (HKCCPN). The split-half reliability and intra-class correlations testified to the high internal reliability of the ratings. Their convergent and discriminant patterns in relations to other psycholinguistic measures echoed previous findings reported on Chinese. There were high correlations for semantic radical transparency, imageability and concreteness, and moderate-to-high correlations for AoA and familiarity among subsets of items that had been collected in previous studies. Concurrent validity analyses showed convergence in predicting behavioral response times in various tasks (lexical decision, naming, and writing-to-dictation) when compared with other Chinese character databases. High predictive validity was shown in writing-to-dictation data from an independent sample of 20 native Cantonese speakers. Several objective psycholinguistic measures (character frequency, stroke number, number of words formed, number of homophones and number of meanings) were included in this database to facilitate its use. These new ratings extend the currently available norms in language and reading research in Cantonese Chinese for researchers, clinicians, and educators, as well as provide them with a wider choice of stimuli.
We introduce a novel dataset of affective, semantic, and descriptive norms for all facial emojis at the point of data collection. We gathered and examined subjective ratings of emojis from 138 German speakers along five essential dimensions: valence, arousal, familiarity, clarity, and visual complexity. Additionally, we provide absolute frequency counts of emoji use, drawn from an extensive Twitter corpus, as well as a much smaller WhatsApp database. Our results replicate the well-established quadratic relationship between arousal and valence of lexical items, also known for words. We also report associations among the variables: for example, the subjective familiarity of an emoji is strongly correlated with its usage frequency, and positively associated with its emotional valence and clarity of meaning. We establish the meanings associated with face emojis, by asking participants for up to three descriptions for each emoji. Using this linguistic data, we computed vector embeddings for each emoji, enabling an exploration of their distribution within the semantic space. Our description-based emoji vector embeddings not only capture typical meaning components of emojis, such as their valence, but also surpass simple definitions and direct emoji2vec models in reflecting the semantic relationship between emojis and words. Our dataset stands out due to its robust reliability and validity. This new semantic norm for face emojis impacts the future design of highly controlled experiments focused on the cognitive processing of emojis, their lexical representation, and their linguistic properties.
In the domain of cognitive studies on the lexico-semantic representational system, one of the most important means of ensuring effective experimental designs is using ecological stimulus sets accompanied by normative data on the most relevant variables affecting the processing of their items. In the context of image sets, color photographs are particularly suited to this purpose as they reduce the difficulty of visual decoding processes that may emerge with traditional image sets of line drawings. This is especially so in clinical populations. In this study we provide Italian norms for a set of 357 high quality image-items belonging to 23 semantic subcategories from the Moreno-Martínez and Montoro database. Data from several variables affecting image processing were collected from a sample of 255 Italian-speaking participants: age of acquisition, familiarity, lexical frequency, manipulability, name agreement, typicality and visual complexity. Lexical frequency data were derived from the CoLFIS corpus. Furthermore, we collected data on image oral naming latencies to explore how the variance in these latencies could be explained by these critical variables. Multiple regression analyses on the naming latencies show classical psycholinguistic phenomena, such as the effects of age of acquisition and name agreement. In addition, manipulability was also a significant predictor. The described Italian normative data and naming latencies are available for download as supplementary material.
Research on language and cognition relies extensively on psycholinguistic datasets or "norms". These datasets contain judgments of lexical properties like concreteness and age of acquisition, and can be used to norm experimental stimuli, discover empirical relationships in the lexicon, and stress-test computational models. However, collecting human judgments at scale is both time-consuming and expensive. This issue of scale is compounded for multi-dimensional norms and those incorporating context. The current work asks whether large language models (LLMs) can be leveraged to augment the creation of large, psycholinguistic datasets in English. I use GPT-4 to collect multiple kinds of semantic judgments (e.g., word similarity, contextualized sensorimotor associations, iconicity) for English words and compare these judgments against the human "gold standard". For each dataset, I find that GPT-4's judgments are positively correlated with human judgments, in some cases rivaling or even exceeding the average inter-annotator agreement displayed by humans. I then identify several ways in which LLM-generated norms differ from human-generated norms systematically. I also perform several "substitution analyses", which demonstrate that replacing human-generated norms with LLM-generated norms in a statistical model does not change the sign of parameter estimates (though in select cases, there are significant changes to their magnitude). I conclude by discussing the considerations and limitations associated with LLM-generated norms in general, including concerns of data contamination, the choice of LLM, external validity, construct validity, and data quality. Additionally, all of GPT-4's judgments (over 30,000 in total) are made available online for further analysis.
The aim of this research is to present a Spanish Word Association Norms (WAN) database of concrete nouns. The database includes 234 stimulus words (SWs) and 67,622 response words (RWs) provided by 478 young Mexican adults. Eight different measures were calculated to quantitatively analyze word-word relationships: 1) Associative strength of the first associate, 2) Associative strength of the second associate, 3) Sum of associative strength of first two associates, 4) Difference in associative strength between first two associates, 5) Number of different associates, 6) Blank responses, 7) Idiosyncratic responses, and 8) Cue validity of the first associate. The resulting database is an important contribution given that there are no published word association norms for Mexican Spanish. The results of this study are an important resource for future research regarding lexical networks, priming effects, semantic memory, among others.
Age of acquisition (AoA) is an important psycholinguistic variable that affects the performance of healthy individuals and patients in a large variety of cognitive tasks. For this reason, it becomes more and more compelling to collect new AoA norms for a large set of stimuli in order to allow better control and manipulation of AoA in future research. An important motivation of the present study is to extend previous Italian norms by collecting AoA ratings for a much larger range of Italian words for which concreteness and semantic-affective norms are now available thus ensuring greater coverage of words varying along these dimensions. In the present study, we collected AoA ratings for 1,957 Italian content words (adjectives, nouns, and verbs), by asking healthy adult participants to estimate the age at which they thought they had learned the word in a Web survey procedure. First, we found high split-half correlation within our sample, suggesting strong internal reliability. Second, our data indicate that the ratings collected in this study are as valid and reliable as those collected in previous studies for Italian across different age populations (adult and children) and other languages. Finally, we analyzed the relation between AoA ratings and other lexical-semantic variables (e.g., word frequency, imageability, valence, arousal) and showed that these correlations were generally consistent with the correlations reported in other normative studies for Italian and other languages. Therefore, our new AoA norms are a valuable source of information for future research in the Italian language. The full database is available at the Open Science Framework (osf.io/3trg2).
This study investigated the lexical-semantic space organized by the semantic and affective features of Indonesian words and their relationship with gender and cultural aspects. We recruited 1,402 participants who were native speakers of Indonesian to rate affective and lexico-semantic properties of 1,490 Indonesian words. Valence, Arousal, Dominance, Predictability, Subjective Frequency, and Concreteness ratings were collected for each word from at least 52 people. We explored cultural differences between American English ANEW (affective norms for English words), Spanish ANEW, and the new Indonesian inventory [called CEFI (concreteness, emotion, and subjective frequency norms for Indonesian words)]. We found functional relationships between the affective dimensions that were similar across languages, but also cultural differences dependent on gender.
This paper introduces a novel collection of word embeddings, numerical representations of lexical semantics, in 55 languages, trained on a large corpus of pseudo-conversational speech transcriptions from television shows and movies. The embeddings were trained on the OpenSubtitles corpus using the fastText implementation of the skipgram algorithm. Performance comparable with (and in some cases exceeding) embeddings trained on non-conversational (Wikipedia) text is reported on standard benchmark evaluation datasets. A novel evaluation method of particular relevance to psycholinguists is also introduced: prediction of experimental lexical norms in multiple languages. The models, as well as code for reproducing the models and all analyses reported in this paper (implemented as a user-friendly Python package), are freely available at: https://github.com/jvparidon/subs2vec.
This article presents CPB-LEX, a large-scale database of lexical statistics derived from children's picture books (age range 0-8 years). Such a database is essential for research in psychology, education and computational modelling, where rich details on the vocabulary of early print exposure are required. CPB-LEX was built through an innovative method of computationally extracting lexical information from automatic speech-to-text captions and subtitle tracks generated from social media channels dedicated to reading picture books aloud. It consists of approximately 25,585 types (wordforms) and their frequency norms (raw and Zipf-transformed), a lexicon of bigrams (two-word sequences and their transitional probabilities) and a document-term matrix (which shows the importance of each word in the corpus in each book). Several immediate contributions of CPB-LEX to behavioural science research are reported, including that the new CPB-LEX frequency norms strongly predict age of acquisition and outperform comparable child-input lexical databases. The database allows researchers and practitioners to extract lexical statistics for high-frequency words which can be used to develop word lists. The paper concludes with an investigation of how CPB-LEX can be used to extend recent modelling research on the lexical diversity children receive from picture books in addition to child-directed speech. Our model shows that the vocabulary input from a relatively small number of picture books can dramatically enrich vocabulary exposure from child-directed speech and potentially assist children with vocabulary input deficits. The database is freely available from the Open Science Framework repository: https://tinyurl.com/4este73c.
In this study, we present the first database of pictures and their corresponding psycholinguistic norms for Polish: the CLT database. In this norming study, we used the pictures from Cross-Linguistic Lexical Tasks (CLT): a set of colored drawings of 168 object and 146 actions. The CLT pictures were carefully created to provide a valid tool for multicultural comparisons. The pictures are accompanied by norms for Naming latencies, Name agreement, Goodness of depiction, Image agreement, Concept familiarity, Age of acquisition, Imageability, Lexical frequency, and Word complexity. We also report analyses of predictors of Naming latencies for pictures of objects and actions. Our results show that Name agreement, Concept familiarity, and Lexical frequency are significant predictors of Naming latencies for pictures of both objects and actions. Additionally, Age of acquisition significantly predicts Naming latencies of pictures of objects. The CLT database is freely available at osf.io/gp9qd. The full set of CLT pictures, including additional variants of pictures, is available on request at osf.io/y2cwr.
Modality exclusivity norms have been developed in different languages for research on the relationship between perceptual and conceptual systems. This paper sets up the first modality exclusivity norms for Chinese, a Sino-Tibetan language with semantics as its orthographically relevant level. The norms are collected through two studies based on Chinese sensory words. The experimental designs take into consideration the morpho-lexical and orthographic structures of Chinese. Study 1 provides a set of norms for Mandarin Chinese single-morpheme words in mean ratings of the extent to which a word is experienced through the five sense modalities. The degrees of modality exclusivity are also provided. The collected norms are further analyzed to examine how sub-lexical orthographic representations of sense modalities in Chinese characters affect speakers' interpretation of the sensory words. In particular, we found higher modality exclusivity rating for the sense modality explicitly represented by a semantic radical component, as well as higher auditory dominant modality rating for characters with transparent phonetic symbol components. Study 2 presents the mean ratings and modality exclusivity of coordinate disyllabic compounds involving multiple sense modalities. These studies open new perspectives in the study of modality exclusivity. First, links between modality exclusivity and writing systems have been established which has strengthened previous accounts of the influence of orthography in the processing of visual information in reading. Second, a new set of modality exclusivity norms of compounds is proposed to show the competition of influence on modality exclusivity from different linguistic factors and potentially allow such norms to be linked to studies on synesthesia and semantic transparency.
Iconicity, understood as a resemblance relationship between meaning and form, is an important variable that has important psycholinguistic effects in lexical processing and language learning across modalities of language. With the growing interest in iconicity, clear operationalizations in terms of the different ways in which iconicity is construed and measured are critical for establishing its broader psycholinguistic profile. This study reports a normed database of iconicity ratings for the same concepts in British Sign Language (BSL) and German Sign Language (DGS). As a related dimension, we also report the type of iconic mapping strategy, i.e., a nominal variable that reflects the different ways in which signs make form-meaning associations for each sign. Finally, we include concreteness ratings for the same concepts. Data from deaf and hearing signers show that iconicity ratings are strongly correlated across both languages, with different distributions across the different strategies, and skewed towards the iconic end of the scale for all groups except German hearing non-signers. Concreteness ratings in BSL and DGS are correlated, though more weakly, and skewed towards the concrete end of the scale. Interestingly, this differs from findings for spoken languages, where concreteness ratings exhibit substantially stronger correlations and abstract concepts are more predominantly represented. We also find that iconicity and concreteness ratings have a moderate positive and strong positive correlation in BSL and DGS, respectively. These results will be useful in psycholinguistic research and highlight differences that can be attributed to the manual-visual modality of signs.
Abstract Sign language offers a unique perspective on the human faculty of language by illustrating that linguistic abilities are not bound to speech and writing. In studies of spoken and written language processing, lexical variables such as, for example, age of acquisition have been found to play an important role, but such information is not as yet available for German Sign Language ( Deutsche Gebärdensprache, DGS). Here, we present a set of norms for frequency, age of acquisition, and iconicity for more than 300 lexical DGS signs, derived from subjective ratings by 32 deaf signers. We also provide additional norms for iconicity and transparency for the same set of signs derived from ratings by 30 hearing non-signers. In addition to empirical norming data, the dataset includes machine-readable information about a sign’s correspondence in German and English, as well as annotations of lexico-semantic and phonological properties: one-handed vs. two-handed, place of articulation, most likely lexical class, animacy, verb type, (potential) homonymy, and potential dialectal variation. Finally, we include information about sign onset and offset for all stimulus clips from automated motion-tracking data. All norms, stimulus clips, data, as well as code used for analysis are made available through the Open Science Framework in the hope that they may prove to be useful to other researchers: 10.17605/OSF.IO/MZ8J4
Human ratings of valence, arousal, and dominance are frequently used to study the cognitive mechanisms of emotional attention, word recognition, and numerous other phenomena in which emotions are hypothesized to play an important role. Collecting such norms from human raters is expensive and time consuming. As a result, affective norms are available for only a small number of English words, are not available for proper nouns in English, and are sparse in other languages. This paper investigated whether affective ratings can be predicted from length, contextual diversity, co-occurrences with words of known valence, and orthographic similarity to words of known valence, providing an algorithm for estimating affective ratings for larger and different datasets. Our bootstrapped ratings achieved correlations with human ratings on valence, arousal, and dominance that are on par with previously reported correlations across gender, age, education and language boundaries. We release these bootstrapped norms for 23,495 English words.
Project on Linguistic Analysis, Berkeley.
This paper gives a brief survey of the Saarbrücken project on Old Icelandic legal texts, sponsored by the German Research Society, within the Special Research Area “Computer linguistics.” The project's main points of interest are (1) producing adequate machine-readable versions and parsed indices of all legal texts in Old Icelandic, (2) graphemic studies of legal manuscripts, and (3) studies of the distribution and valence of the verbs in those texts. A proposal for encoding Old Norse/Old Icelandic demonstrates how texts of different standards (normalized, diplomatic, graphetic) can be encoded as compatibly as possible. The description of a combined normalization-lemmatization process reveals that even little normalization in a diplomatic text will save much manual parsing.
The Opera del Vocabolario Italiano was given a mandate in 1964 to create a Historical Dictionary of the Italian Language. The main objective was to provide a tool which would give vital information on the development of the Italian language from its origins to the present day. In 1986 the Center incorporated modern computer technology into the project and this led to a series of decisions which affected the nature and the outcome of the project. This article traces the development of the project, and describes both hardware and software systems used, as well as the nature of the relational database being created and its linguistic applications.
The Corpus dei Manoscritti Copti Letterari is a project whose original aim was to reconstruct the Coptic codices from the White Monastery in Upper Egypt. The project was later expanded to include all Coptic literature. In 1980 a new project was launched to transfer the data into machine-readable form and make the information available, in as generic a format as possible, to scholars throughout the world.
CLIPON is an acronym for Concordanze della Lingua Italiana Poetica dell'Otto/Novecento. The aim of the project described here is to produce lexicons and lemmatized concordances of the literary Italian language of the nineteenth and twentieth centuries. The corpus involves groups of mainly poetic works and authors that have a common denominator as regards schools, currents, culture and chronology.
Statistical information on a substantial corpus of representative Spanish texts is needed in order to determine the significance of data about individual authors or texts by means of comparison. This study describes the organization and analysis of a 150,000-word corpus of 30 well-known twentieth-century Spanish authors. Tables show the computational results of analyses involving sentences, segments, quotations, and word length.
This article summarizes the activities of the Istituto di Linguistica Computazionale. We discuss the Italian Multi-functional Lexical Databases; the projects focussing on linguistic analysis and generation; corpora in the MRF, textual databases and linguistic workstations; computer-assisted humanities teaching; and the various cooperative ventures, seminars and conferences offered by the Institute.
The Century of Prose Corpus is a historical corpus of British English of the period 1680–1780. It has been designed to provide a resource for students of the language of that era. The COPC is diachronic and may be considered a unit in what will eventually become a series of corpora providing access to the whole of the English language from the oldest specimens to the present. This article describes and explains the various features of the COPC.
This paper concerns the Charrette Project, a multimedia electronic archive of a medieval manuscript tradition. In this paper, we argue that the computer's strengths in manipulating complex and varied resources should be an important organizing principle in the conception and construction of electronic text projects. Specifically, we describe the elements of the Charrette archive, its architecture, and its potential for scholarly research and pedagogical applications.
This article is a detailed account of COMLEX Syntax, an on-line syntactic dictionary of English, developed by the Proteus Project at New York University under the auspices of the Linguistics Data Consortium. This lexicon was intended to be used for a variety of tasks in natural language processing by computer and as such has very detailed classes with a large number of syntactic features and complements for the major parts of speech and is, as far as possible, theory neutral. The dictionary was entered by hand with reference to hard copy dictionaries, an on-line concordance and native speakers‘intuition. Thus it is without prior encumbrances and can be used for both pure research and commercial purposes.
In this paper, we study the problem of adding a large number of new words into a Chinese thesaurus according to their definitions in a Chinese dictionary, while minimizing the effort of hand tagging. To deal with the problem, we first make use of a kind of supervised learning technique to learn a set of defining formats for each class in the thesaurus, which tries to characterize the regularities about the definitions of the words in the class. We then use traditional techniques in Graph theory to derive a minimal subset of the new words to be added into the thesaurus, which meets the following condition: if we add the new words in the subset into the thesaurus by hand, the other new words can be added into the thesaurus automatically by matching their definitions with the defining formats of each class in the thesaurus. The method uses little, if any, language-specific or thesaurus-specific knowledge, and can be applied to the thesauri of other languages.
This paper discusses the design of the EuroWordNet database, in which semantic databases like WordNet1.5 for several languages are combined via a so-called inter-lingual-index. In this database, language-independent data is shared whilst language-specific properties are maintained. A special interface has been developed to compare the semantic configurations across languages and to track down differences.
We discuss ways in which EuroWordNet (EWN) can be used in multilingual information retrieval activities, focusing on two approaches to Cross-Language Text Retrieval that use the EWN database as a large-scale multilingual semantic resource. The first approach indexes documents and queries in terms of the EuroWordNet Inter-Lingual-Index, thus turning term weighting and query/document matching into language-independent tasks. The second describes how the information in the EWN database could be integrated with a corpus-based technique, thus allowing retrieval of domain-specific terms that may not be present in our multilingual database. Our objective is to show the potential of EuroWordNet as a promising alternative to existing approaches to Cross-Language Text Retrieval.