1396 norm sets
Large-scale word association datasets are both important tools used in psycholinguistics and used as models that capture meaning when considered as semantic networks. Here, we present word association norms for Rioplatense Spanish, a variant spoken in Argentina and Uruguay. The norms were derived through a large-scale crowd-sourced continued word association task in which participants give three associations to a list of cue words. Covering over 13,000 words and +3.6 M responses, it is currently the most extensive dataset available for Spanish. We compare the obtained dataset with previous studies in Dutch and English to investigate the role of grammatical gender and studies that used Iberian Spanish to test generalizability to other Spanish variants. Finally, we evaluated the validity of our data in word processing (lexical decision reaction times) and semantic (similarity judgment) tasks. Our results demonstrate that network measures such as in-degree provide a good prediction of lexical decision response times. Analyzing semantic similarity judgments showed that results replicate and extend previous findings demonstrating that semantic similarity derived using spreading activation or spectral methods outperform word embeddings trained on text corpora.
Sensorimotor information is vital to the conceptual representation of our knowledge system. This study collects perceptual and action ratings for 664 disyllabic nouns among 438 native speakers and creates the first and largest dataset of sensorimotor norms for nouns in Chinese. Using aggregated semantic covariates, including concreteness ratings from a concreteness rating study, as well as the reaction times and error rates from a lexical decision study, our current work demonstrates the strengths of sensory modalities and action effectors in Chinese nouns and explores the contributions of embodied experiences in reflecting orthographic representations and semantic processing in the Chinese language. This study contributes valuable data sources to the study of Chinese lexical processing and highlights the importance of sensorimotor information and embodied manifestations in the semantic representations of concepts. Our results also support the language universal that orthographic awareness in lexical processing and reading supersedes phonological awareness.
We introduce a novel dataset of affective, semantic, and descriptive norms for all facial emojis at the point of data collection. We gathered and examined subjective ratings of emojis from 138 German speakers along five essential dimensions: valence, arousal, familiarity, clarity, and visual complexity. Additionally, we provide absolute frequency counts of emoji use, drawn from an extensive Twitter corpus, as well as a much smaller WhatsApp database. Our results replicate the well-established quadratic relationship between arousal and valence of lexical items, also known for words. We also report associations among the variables: for example, the subjective familiarity of an emoji is strongly correlated with its usage frequency, and positively associated with its emotional valence and clarity of meaning. We establish the meanings associated with face emojis, by asking participants for up to three descriptions for each emoji. Using this linguistic data, we computed vector embeddings for each emoji, enabling an exploration of their distribution within the semantic space. Our description-based emoji vector embeddings not only capture typical meaning components of emojis, such as their valence, but also surpass simple definitions and direct emoji2vec models in reflecting the semantic relationship between emojis and words. Our dataset stands out due to its robust reliability and validity. This new semantic norm for face emojis impacts the future design of highly controlled experiments focused on the cognitive processing of emojis, their lexical representation, and their linguistic properties.
Cue reactivity is essential to the maintenance of addictive disorders. A useful way to study cue reactivity is by means of normative pictures, but few validated tobacco-related pictures are available. This study describes a database of smoking-related pictures: The Geneva Smoking Pictures (GSP). Sixty smoking-related pictures were presented to 91 participants who assessed them according to the classic emotional pictures validation provided by the International Affective Picture System (NIMH Center for the Study of Emotion and Attention, 2002). The pictures were rated according to three dimensions: (1) valence (from positive to negative), (2) emotional arousal (from high arousing to low arousing), and (3) dominance (from submissive to dominant). Participants were also screened with the Fagerström Test for Nicotine Dependence. Normative ratings for valence, arousal and dominance of the pictures are provided for the whole sample, as well as separately for dependent (n = 46) and nondependent smokers (n = 45). Arousal and dominance were associated with greater nicotine dependence, but valence ratings were not. The GSP is a normative database providing a large number of stimuli for investigators who are conducting nicotine and tobacco research.
Age of acquisition (AoA) is an important psycholinguistic variable that affects the performance of healthy individuals and patients in a large variety of cognitive tasks. For this reason, it becomes more and more compelling to collect new AoA norms for a large set of stimuli in order to allow better control and manipulation of AoA in future research. An important motivation of the present study is to extend previous Italian norms by collecting AoA ratings for a much larger range of Italian words for which concreteness and semantic-affective norms are now available thus ensuring greater coverage of words varying along these dimensions. In the present study, we collected AoA ratings for 1,957 Italian content words (adjectives, nouns, and verbs), by asking healthy adult participants to estimate the age at which they thought they had learned the word in a Web survey procedure. First, we found high split-half correlation within our sample, suggesting strong internal reliability. Second, our data indicate that the ratings collected in this study are as valid and reliable as those collected in previous studies for Italian across different age populations (adult and children) and other languages. Finally, we analyzed the relation between AoA ratings and other lexical-semantic variables (e.g., word frequency, imageability, valence, arousal) and showed that these correlations were generally consistent with the correlations reported in other normative studies for Italian and other languages. Therefore, our new AoA norms are a valuable source of information for future research in the Italian language. The full database is available at the Open Science Framework (osf.io/3trg2).
This paper introduces association norms of German noun compounds as a lexical-semantic resource for cognitive and computational linguistics research on compositionality. Based on an existing database of German noun compounds, we collected human associations to the compounds and their constituents within a web experiment. The current study describes the collection process and a part-of-speech analysis of the association resource. In addition, we demonstrate that the associations provide insight into the semantic properties of the compounds, and perform a case study that predicts the degree of compositionality of the experiment compound nouns, as relying on the norms. Applying a comparatively simple measure of association overlap, we reach a Spearman rank correlation coefficient of rs = 0.5228, p <.000001, when comparing our predictions with human judgements.
Iconicity, understood as a resemblance relationship between meaning and form, is an important variable that has important psycholinguistic effects in lexical processing and language learning across modalities of language. With the growing interest in iconicity, clear operationalizations in terms of the different ways in which iconicity is construed and measured are critical for establishing its broader psycholinguistic profile. This study reports a normed database of iconicity ratings for the same concepts in British Sign Language (BSL) and German Sign Language (DGS). As a related dimension, we also report the type of iconic mapping strategy, i.e., a nominal variable that reflects the different ways in which signs make form-meaning associations for each sign. Finally, we include concreteness ratings for the same concepts. Data from deaf and hearing signers show that iconicity ratings are strongly correlated across both languages, with different distributions across the different strategies, and skewed towards the iconic end of the scale for all groups except German hearing non-signers. Concreteness ratings in BSL and DGS are correlated, though more weakly, and skewed towards the concrete end of the scale. Interestingly, this differs from findings for spoken languages, where concreteness ratings exhibit substantially stronger correlations and abstract concepts are more predominantly represented. We also find that iconicity and concreteness ratings have a moderate positive and strong positive correlation in BSL and DGS, respectively. These results will be useful in psycholinguistic research and highlight differences that can be attributed to the manual-visual modality of signs.
The present study introduces the Extreme Climate Event Database (EXCEED), a picture database intended to induce emotionally salient stimuli reactions in the context of natural hazards associated with global climate change and related extreme events. The creation of the database was motivated by the need to better understand the impact that the increase in natural disasters worldwide has on human emotional reactions. This new database consists of 150 pictures divided into three categories: two negative categories that depict images of floods and droughts, and a neutral category composed of inanimate objects. Affective ratings were obtained using online survey software from 50 healthy Brazilian volunteers who rated the pictures according to valence and arousal, which are two fundamental dimensions used to describe emotional experiences. Valence refers to the appraisal of pleasantness conveyed by a stimulus, and arousal involves internal emotional activation induced by a stimulus. Data from picture rating, sex difference in affective ratings and psychometric properties of the database are presented here. Together, the data validate the use of EXCEED in research related to natural hazards and human reactions.
This article presents CPB-LEX, a large-scale database of lexical statistics derived from children's picture books (age range 0-8 years). Such a database is essential for research in psychology, education and computational modelling, where rich details on the vocabulary of early print exposure are required. CPB-LEX was built through an innovative method of computationally extracting lexical information from automatic speech-to-text captions and subtitle tracks generated from social media channels dedicated to reading picture books aloud. It consists of approximately 25,585 types (wordforms) and their frequency norms (raw and Zipf-transformed), a lexicon of bigrams (two-word sequences and their transitional probabilities) and a document-term matrix (which shows the importance of each word in the corpus in each book). Several immediate contributions of CPB-LEX to behavioural science research are reported, including that the new CPB-LEX frequency norms strongly predict age of acquisition and outperform comparable child-input lexical databases. The database allows researchers and practitioners to extract lexical statistics for high-frequency words which can be used to develop word lists. The paper concludes with an investigation of how CPB-LEX can be used to extend recent modelling research on the lexical diversity children receive from picture books in addition to child-directed speech. Our model shows that the vocabulary input from a relatively small number of picture books can dramatically enrich vocabulary exposure from child-directed speech and potentially assist children with vocabulary input deficits. The database is freely available from the Open Science Framework repository: https://tinyurl.com/4este73c.
In this study, we present the first database of pictures and their corresponding psycholinguistic norms for Polish: the CLT database. In this norming study, we used the pictures from Cross-Linguistic Lexical Tasks (CLT): a set of colored drawings of 168 object and 146 actions. The CLT pictures were carefully created to provide a valid tool for multicultural comparisons. The pictures are accompanied by norms for Naming latencies, Name agreement, Goodness of depiction, Image agreement, Concept familiarity, Age of acquisition, Imageability, Lexical frequency, and Word complexity. We also report analyses of predictors of Naming latencies for pictures of objects and actions. Our results show that Name agreement, Concept familiarity, and Lexical frequency are significant predictors of Naming latencies for pictures of both objects and actions. Additionally, Age of acquisition significantly predicts Naming latencies of pictures of objects. The CLT database is freely available at osf.io/gp9qd. The full set of CLT pictures, including additional variants of pictures, is available on request at osf.io/y2cwr.
Modality exclusivity norms have been developed in different languages for research on the relationship between perceptual and conceptual systems. This paper sets up the first modality exclusivity norms for Chinese, a Sino-Tibetan language with semantics as its orthographically relevant level. The norms are collected through two studies based on Chinese sensory words. The experimental designs take into consideration the morpho-lexical and orthographic structures of Chinese. Study 1 provides a set of norms for Mandarin Chinese single-morpheme words in mean ratings of the extent to which a word is experienced through the five sense modalities. The degrees of modality exclusivity are also provided. The collected norms are further analyzed to examine how sub-lexical orthographic representations of sense modalities in Chinese characters affect speakers' interpretation of the sensory words. In particular, we found higher modality exclusivity rating for the sense modality explicitly represented by a semantic radical component, as well as higher auditory dominant modality rating for characters with transparent phonetic symbol components. Study 2 presents the mean ratings and modality exclusivity of coordinate disyllabic compounds involving multiple sense modalities. These studies open new perspectives in the study of modality exclusivity. First, links between modality exclusivity and writing systems have been established which has strengthened previous accounts of the influence of orthography in the processing of visual information in reading. Second, a new set of modality exclusivity norms of compounds is proposed to show the competition of influence on modality exclusivity from different linguistic factors and potentially allow such norms to be linked to studies on synesthesia and semantic transparency.
L’objectif premier de ce travail etait de caracteriser les images proposees par Bonin, Peereman, Malardier, Meot et Chalard (2003) en termes d’âge d’acquisition (AoA) objectif, recueilli aupres d’enfants âges de 2: 6 a 10: 11 ans. (http:// www. unice. fr/ LPEQ/ base_ AoA/ aoa_ intro. php). La comparaison avec les normes en francais des images de Snodgrass et Vanderwart (1980) montre que les nouvelles images correspondent, en moyenne, a des mots plus rares et d’AoA plus tardif. Cependant, correlations et regressions multiples indiquent que les memes facteurs sont impliques dans l’emergence de l’AoA (variabilite d’imagerie et frequence lexicale) a travers les deux jeux d’images. De plus, quelle que soit la base d’images, l’AoA est significativement mieux correle et specifiquement predit par les frequences lexicales du vocabulaire de l’enfant (NOVLEX, MANULEX) que par les frequences calculees sur des corpus du vocabulaire de l’adulte (BRULEX, LEXIQUE).
This study is a cross-linguistic, conceptual replication of Lynott and Connell’s (2009, 2013) modality exclusivity norms. Their English properties and concepts were translated into Dutch, then independently tested as follows. Forty-two respondents rated the auditory, haptic, and visual strength of those words. Mean ratings were then computed, with a high interrater reliability and interitem consistency. Based on the three modalities, each word also features a specific modality exclusivity, and a dominant modality. The norms also include external measures of word frequency, length, distinctiveness, age of acquisition, and known percentage. Starting with the results, unimodal, bimodal, and tri-modal words appear. Visual and haptic experience are quite related, leaving a more independent auditory experience. These different relations are important because they may correlate with different levels of detail in word comprehension (Louwerse & Connell, 2011). Auditory and visual words tend towards unimodality, whereas haptic words tend towards multimodality. Likewise, properties are more unimodal than concepts. The form of words is not quite as arbitrary as we used to think. It is connected to their meaning. This 'sound symbolism' was tested by means of a regression: Auditory strength predicts lexical properties of the words (frequency, distinctiveness...) better than the other modalities do, or else with a different polarity. Last, words from these norms were used as the stimuli for an experiment, in which switches across modalities incurred processing costs (Bernabeu, Willems, & Louwerse, 2017). - Dashboard for using Dutch modality norms (336 properties, 411 concepts) and exploring various analyses with them.<br> - A summary may be found here. - The entire data set and analysis code are available.<strong><br> </strong> References Bernabeu, P., Willems, R. M., & Louwerse, M. M. (2017). Modality switch effects emerge early and increase throughout conceptual processing: Evidence from ERPs. In G. Gunzelmann, A. Howes, T. Tenbrink, & E. J. Davelaar (Eds.), <em>Proceedings of the 39th Annual Conference of the Cognitive Science Society</em> (pp. 1629-1634). Austin, TX: Cognitive Science Society. Louwerse, M., & Connell, L. (2011). A taste of words: linguistic context and perceptual simulation predict the modality of words. <em>Cognitive Science, 35, </em>2, 381-98. Lynott, D., & Connell, L. (2009). Modality exclusivity norms for 423 object properties. <em>Behavior Research Methods, 41, </em>2, 558-564. Lynott, D., & Connell, L. (2013). Modality exclusivity norms for 400 nouns: The relationship between perceptual experience and surface word form.<em> Behavior Research Methods, 45</em>, 516-526.
Many abstract words refer to internal cognitive events or states, such as thinking or believing, or to cognitive products, such as theories, ideas, or whims (Binder et al., Cognitive Neuropsychology, 33, 130–174, 2016). Mental state information is proposed to be an important component in the grounding of abstract meaning (Kiefer et al., 2022, Muraki et al., 2022), such that our inner cognitive experiences form a foundational aspect of semantic representation. We tested this proposal by first collecting cognition ratings for over 8000 English words. Then, we used the norms generated from our ratings to examine the unique variance explained by cognition ratings in performance on lexical-semantic tasks. We found a significant effect of cognition, such that there was a facilitative relationship between cognition ratings and behavioral responses, even when controlling for other key lexical and semantic variables. Specifically, words rated as more cognitive in nature elicited faster and more accurate task responses, especially for words with more abstract meanings. This study highlights a novel behavioral effect that is consistent with a multidimensional account of semantic representation.
This work presents a lexical database with cognate annotation and phonological alignment for over 6,500 documented language varieties. The database includes per-family and global phylogenetic resources and offers a pre-computed global tree for language variety distance from normalized trees obtained with Bayesian Markov Chain Monte Carlo (MCMC) inference. Lexical data is provided in a single tabular file for convenience of usage, and resources are built adhering to best practices and state-of-the-art algorithms for historical linguistics. The database is a convenient source for research prototypes, method development, and analysis bootstrap. All resources are freely available for download for all interested researchers.
We introduce JAMBU, a cognate database of South Asian languages which unifies dozens of previous sources in a structured and accessible format. The database includes nearly 287k lemmata from 602 lects, grouped together in 23k sets of cognates. We outline the data wrangling necessary to compile the dataset and train neural models for reflex prediction on the Indo- Aryan subset of the data. We hope that JAMBU is an invaluable resource for all historical linguists and Indologists, and look towards further improvement and expansion of the database.
コーパス言語学の歴史は長く、情報革命に伴って、次第に大きな展開をとげることになった。本論文では、先ず、コーパスの構築とその適用に関する諸問題を見て、次に、デジタル映像を取り込んだコーパスの作成を提案する。最近の技術の発展により、語彙のデータベースは比較的簡単に映像情報とリンクさせることができ、高度なコンコーダンスを作成することが可能となった。これにより、学習者は、調べたい語彙に関して、その細かいニュアンスや、実際に使われている場面等を知ることができ、語用言語学的間違いを防ぐことができるようになる。さらに、本論文では、そのようなコーパスの作成の難しさを著作権等の問題から論じ、最後に、実例を紹介する。
The goal of this paper is to describe how adjectives are encoded in Cornetto, a semantic lexical database for Dutch. Cornetto combines two existing lexical resources with different semantic organisation, i.e. Dutch Wordnet (DWN) with a synset organisation and Referentie Bestand Nederlands (RBN) with an organisation in Lexical Units. Both resources will be aligned and mapped on the formal ontology SUMO. In this paper, we will first present details of the description of adjectives in each of the the two resources. We will then address the problems that are encountered during alignment to the SUMO ontology which are greatly due to the fact that SUMO has never been tested for its adequacy with respect to adjectives. We contrasted SUMO with an existing semantic classification which resulted in a further refined and extended SUMO geared for the description of adjectives.
We propose a new method for empirically determining lists of basic concepts for the purpose of compiling extensive lexicostatistical databases. The idea is to approximate a notion of “swadeshness” formally and reproducibly without expert knowledge or bias, and being able to rank any number of concepts given enough data. Unlike previous approaches, our procedure indirectly measures both stability of concepts against lexical replacement, and their proneness to phenomena such as onomatopoesia and extensive borrowing. The method provides a fully automated way to generate customized Swadesh lists of any desired length, possibly adapted to a given geographical region. We apply the method to a large lexical database of Northern Eurasia, deriving a swadeshness ranking for more than 5,000 concepts expressed by German lemmas. We evaluate this ranking against existing shorter lists of basic concepts to validate the method, and give an English version of the 300 top concepts according to this ranking.
This study established psycholinguistic norms on Cantonese conceptual semantic features (CanFeat) with 954 concepts and semantic features, including feature frequency, sharing degree, and type of each feature. We provided concept-level variables, including participant count, feature counts (total, distinct, concept-specific, concept-shared), feature proportions, coverage, category labels (superordinate category and subordinate category), familiarity, and concreteness. We validated the reliability, specificity, and categorical structural consistency of the dataset through multiple analyses, including intra-coder reliability checks, gender comparability assessments, semantic category distinctions, and feature-level comparisons with established norms. We also compared CanFeat with norms in other languages, highlighting both similarities and differences. The present norming offers a timely and valuable resource for researchers in psycholinguistics, particularly those focusing on Cantonese production and semantic representation.
This normative dataset provides perceptual strength ratings for 5,500 Spanish words across five sensory modalities: touch, hearing, sight, smell, and taste. The word set was originally compiled from Spanish lexical databases and used in Díez-Álamo et al. (2019), ensuring broad coverage of psycholinguistic indices. A total of 671 participants rated approximately 200 words each, indicating the extent to which each sensory modality is involved when experiencing the concept (or property, in the case of adjectives), using a scale from 0 (not at all) to 5 (greatly). The procedure follows the methodology established by Lynott and Connell (2009, 2013). The dataset includes raw participant-level data, with fields for participant ID, task number, age, gender, word familiarity, and modality-specific ratings. This resource is valuable for research on conceptual representation, semantic processing, and multimodal lexical analysis in Spanish.
We present an English lexical database which is fuller, more accurate and more consistent than any other. We believe this to be so because the project has been well-planned, with a 12-month intensive planning phase prior to the lexicography beginning; well-resourced, employing a team of fifteen highly experienced lexicographers for a thirty-month main phase; it has had access to the latest corpus and dictionary-editing technology; it has not been constrained to meet any goals other than an accurate description of the language; and it has been led by a team with singular experience in delivering high-quality and innovative resources. The lexicon will be complete in Summer 2010 and will be available for NLP groups, on terms designed to encourage its research use. 1
In this paper, we introduce a set of resources that we have derived from the EST RÉPUBLICAIN CORPUS, a large, freely-available collection of regional newspaper articles in French, totaling 150 million words. Our resources are the result of a full NLP treatment of the EST RÉPUBLICAIN CORPUS: handling of multi-word expressions, lemmatization, part-of-speech tagging, and syntactic parsing. Processing of the corpus is carried out using statistical machine-learning approaches- joint model of data driven lemmatization and partof-speech tagging, PCFG-LA and dependency based models for parsing- that have been shown to achieve state-of-the-art performance when evaluated on the French Treebank. Our derived resources are made freely available, and released according to the original Creative Common license for the EST RÉPUBLICAIN CORPUS. We additionally provide an overview of the use of these resources in various applications, in particular the use of generated word clusters from the corpus to alleviate lexical data sparseness for statistical parsing.
The purpose of this study was to establish psycholinguistic norms for 249 action pictures in Cantonese, a language with few norms available. We provide normative data for rated visual complexity, rated age of acquisition, name agreement, word frequency and rated familiarity in this study. Forty participants were recruited to participate in both timed picture naming and rating experiments. The linear mixed effect analysis revealed that familiarity, visual complexity, and name agreement were significant predictors of action naming in Cantonese. However, AoA did not show any significant effect on action naming, which is consistently observed in previous studies of action picture naming in Chinese. The possible explanation for null effect of AoA on naming latency are discussed. This set of psycholinguistic norms in Cantonese could serve as a valuable resource for future psycholinguistic, neurolinguistic and clinical studies in Cantonese.
This paper describes the design and construction of a lexical database for Turkish adjectives. We used a textual corpus of about one million running words that we collected from on-line newspapers and magazines available on the Internet. The lexicon contains syntactic category, semantic category, gradability, and thesaurus information about adjectives as well as selectional restrictions. It supports Natural Language Processing (NLP) applications such as parsing, text generation, natural language understanding, and information retrieval. It has been implemented as a relational database. The process of building the lexicon from the textual corpus has been performed semi-automatically using a series of extraction programs. We also implemented a Graphical User Interface to the lexicon.
Human ratings of valence, arousal, and dominance are frequently used to study the cognitive mechanisms of emotional attention, word recognition, and numerous other phenomena in which emotions are hypothesized to play an important role. Collecting such norms from human raters is expensive and time consuming. As a result, affective norms are available for only a small number of English words, are not available for proper nouns in English, and are sparse in other languages. This paper investigated whether affective ratings can be predicted from length, contextual diversity, co-occurrences with words of known valence, and orthographic similarity to words of known valence, providing an algorithm for estimating affective ratings for larger and different datasets. Our bootstrapped ratings achieved correlations with human ratings on valence, arousal, and dominance that are on par with previously reported correlations across gender, age, education and language boundaries. We release these bootstrapped norms for 23,495 English words.
Abstract Sign language offers a unique perspective on the human faculty of language by illustrating that linguistic abilities are not bound to speech and writing. In studies of spoken and written language processing, lexical variables such as, for example, age of acquisition have been found to play an important role, but such information is not as yet available for German Sign Language ( Deutsche Gebärdensprache, DGS). Here, we present a set of norms for frequency, age of acquisition, and iconicity for more than 300 lexical DGS signs, derived from subjective ratings by 32 deaf signers. We also provide additional norms for iconicity and transparency for the same set of signs derived from ratings by 30 hearing non-signers. In addition to empirical norming data, the dataset includes machine-readable information about a sign’s correspondence in German and English, as well as annotations of lexico-semantic and phonological properties: one-handed vs. two-handed, place of articulation, most likely lexical class, animacy, verb type, (potential) homonymy, and potential dialectal variation. Finally, we include information about sign onset and offset for all stimulus clips from automated motion-tracking data. All norms, stimulus clips, data, as well as code used for analysis are made available through the Open Science Framework in the hope that they may prove to be useful to other researchers: 10.17605/OSF.IO/MZ8J4
Age of acquisition (AoA) is a widely used variable that estimates when a lexical item is first understood. Existing English AoA norms have been highly influential in psycholinguistics, education, language acquisition, speech-language pathology, and natural language processing, but have focused primarily on single words. Little information is availed for multi-word expressions (MWEs), despite their central role in language use, vocabulary acquisition, representation and processing. The current study contributes AoA estimates for 80,586 English MWEs using a large language model, GPT-4.1-mini, fine-tuned on newly collected crowdsourced human ratings. Ratings were obtained from 96 US-based native English speakers via Prolific, yielding 47,163 ratings for ~3,999 MWEs. After reliability screening, 3,667 BLUP-adjusted means were used for LLM fine-tuning and validation. Fine-tuning substantially improved alignment with hold-out human ratings. The standard GPT-4.1-mini output correlated with human estimates at r =.67, whereas the model fine-tuned on 3,000 items reached r =.85. A final model trained on all reliable crowdsourced estimates was estimated AoAs for the full MWE list. Results showed relationships with existing psycholinguistic variables aligned with those for single-word AoAs, including that earlier-acquired MWEs tended to be more familiar, useful, and frequent. The estimates exhibited predictive validity against test-based student vocabulary data and explained additional variance beyond frequency, utility, and familiarity. These findings indicate that fine-tuned LLMs can provide useful large-scale AoA estimates for MWEs when grounded in human ratings. The new resource is available via OSF (https://tinyurl.com/3e828fj8) and an interactive webpage has been developed for users: https://cgg-projects.github.io/MWEs/.
The present study describes the development and validation of a facial expression database comprising five different horizontal face angles in dynamic and static presentations. The database includes twelve expression types portrayed by eight Japanese models. This database was inspired by the dimensional and categorical model of emotions: surprise, fear, sadness, anger with open mouth, anger with closed mouth, disgust with open mouth, disgust with closed mouth, excitement, happiness, relaxation, sleepiness, and neutral (static only). The expressions were validated using emotion classification and Affect Grid rating tasks [Russell, Weiss, & Mendelsohn, 1989. Affect Grid: A single-item scale of pleasure and arousal. Journal of Personality and Social Psychology, 57(3), 493-502]. The results indicate that most of the expressions were recognised as the intended emotions and could systematically represent affective valence and arousal. Furthermore, face angle and facial motion information influenced emotion classification and valence and arousal ratings. Our database will be available online at the following URL. https://www.dh.aist.go.jp/database/face2017/.
Investigation of affective and semantic dimensions of words is essential for studying word processing. In this study, we expanded Tse et al.'s (Behav Res Methods 49:1503-1519, 2017; Behav Res Methods 55:4382-4402, 2023) Chinese Lexicon Project by norming five word dimensions (valence, arousal, familiarity, concreteness, and imageability) for over 25,000 two-character Chinese words presented in traditional script. Through regression models that controlled for other variables, we examined the relationships among these dimensions. We included ambiguity, quantified by the standard deviation of the ratings of a given lexical variable across different raters, as separate variables (e.g., valence ambiguity) to explore their connections with other variables. The intensity-ambiguity relationships (i.e., between normed variables and their ambiguities, like valence with valence ambiguity) were also examined. In these analyses with a large pool of words and controlling for other lexical variables, we replicated the asymmetric U-shaped valence-arousal relationship, which was moderated by valence and arousal ambiguities. We also observed a curvilinear relationship between valence and familiarity and between valence and concreteness. Replicating Brainerd et al.'s (J Exp Psychol Gen 150:1476-1499, 2021; J Mem Lang 121:104286, 2021) quadratic intensity-ambiguity relationships, we found that the ambiguity of valence, arousal, concreteness, and imageability decreases as the value of these variables is extremely low or extremely high, although this was not generalized to familiarity. While concreteness and imageability were strongly correlated, they displayed different relationships with arousal, valence, familiarity, and valence ambiguity, suggesting their distinct conceptual nature. These findings further our understanding of the affective and semantic dimensions of two-character Chinese words. The normed values of all these variables can be accessed via https://osf.io/hwkv7.
In this article, we present the results of a study carried out in Bogota, Colombia with 210 university students from five different universities pertaining to diverse socio-demographic groups. The objective of the study was to establish the lexical category norms. 56 lexical-semantic categories used by Battig and Montague (1969) in their classic study were employed. More than 7800 words were collected and organized by range and mode. There are no other studies on this subject for Colombian or Latin American Spanish. We hope that the results presented here will be used both in psycholinguistic and language therapy studies. The collected data were compared to one of the category norm studies made for European Spanish.
CzEng 0.9: Large Parallel Treebank with Rich Annotation We describe our ongoing efforts in collecting a Czech-English parallel corpus CzEng. The paper provides full details on the current version 0.9 and focuses on its new features: (1) data from new sources were added, most importantly a few hundred electronically available books, technical documentation and also some parallel web pages, (2) the full corpus has been automatically annotated up to the tectogrammatical layer (surface and deep syntactic analysis), (3) sentence segmentation has been refined, and (4) several heuristic filters to improve corpus quality were implemented. In total, we provide a sentence-aligned automatic parallel treebank of about 8.0 million sentences, 93 million English and 82 million Czech words. CzEng 0.9 is freely available for non-commercial research purposes.
The <em>Corpus of the Epigraphy of the Italian Peninsula in the 1st Millennium BCE</em> (CEIPoM) is a linguistic database which covers the Oscan, Umbrian, Old Sabellic, Messapic and Venetic languages, as well as epigraphic Latin up to 100 BCE. The database is hosted on GitHub and Zenodo, and provides manually annotated linguistic information on all levels of language structure, ranging from phonology to syntax. In providing a high-resolution digital dataset for language varieties that have until now been largely restricted to printed reference works, this corpus opens up new avenues for research into this unique ancient linguistic area.
Understanding how conceptual knowledge is grounded in bodily experience, and to what extent machine systems can acquire such knowledge without direct sensorimotor experience, are central questions in both cognitive science and embodied artificial intelligence research. Large-scale normative resources are essential for investigating these questions empirically, yet such resources remain sparse for non-Indo-European languages. We present a novel normative database for 3,000 lexicalized concepts in Mandarin Chinese, comprising 11-dimensional sensorimotor ratings and unidimensional embodiment ratings collected from 378 native Mandarin speakers. The ratings demonstrate high reliability and strong cross-norm validity with existing Chinese resources, each of which covers fewer words and a subset of the 11 sensorimotor dimensions. In a validation study, we tested new variables derived from a theoretically motivated metric, Perceptual Strength of Embodiment (PSE) (Huang et al., 2025), together with seven common composite variables, on lexical decision tasks. The results suggest that PSE-Sensorimotor and Minkowski-3 are the strongest composite predictors of lexical decision performance, capturing the facilitatory effects of sensorimotor information on lexical processing. A further exploratory study showed that sensorimotor ratings are substantially recoverable from purely linguistic representations using simple regression models (mean Spearman r =.62 across dimensions), though recovery varied markedly: visual and auditory dimensions yielded higher correspondence than chemosensory ones. Representational similarity analysis further showed that the relational geometry of the sensorimotor space is also partially recoverable (r =.540), consistent with the view that distributional language use encodes aspects of embodied conceptual structure.
<b>Dataset for the paper "Implicit Consequentiality Bias in English: A Corpus of 300+ Verbs", appearing in Behaviour Research Methods</b><b><br></b><b>Abstract of Paper</b><b><br></b>This study provides implicit verb consequentiality norms for a corpus of 305 English verbs, for which Ferstl et al. (BRM, 2011) previously provided implicit causality norms. An on-line sentence completion study was conducted, with data analyzed from 124 respondents who completed fragments such as “John liked Mary and so…”. The resulting bias scores are presented in an Appendix, with more detail in supplementary material in the University of Sussex Research Data Repository (via 10.25377/sussex.c.5082122), where we also present lexical and semantic verb features: frequency, semantic class and emotional valence of the verbs. We compare our results with those of our study of implicit causality and with the few published studies of implicit consequentiality. As in our previous study, we also considered effects of gender and verb valence, which requires stable norms for a large number of verbs. The corpus will facilitate future studies in a range of areas, including psycholinguistics and social psychology, particularly those requiring parallel sentence completion norms for both causality and consequentiality.<b><br></b>
The current study presents ratings by 540 Spanish native speakers for dominance, familiarity, subjective age of acquisition (AoA), and sensory experience (SER) for the 875 Spanish words included in the Madrid Affective Database for Spanish (MADS). The norms can be downloaded as supplementary materials for this manuscript from https://figshare.com/s/8e7b445b729527262c88 These ratings may be of potential relevance to researches who are interested in characterizing the interplay between language and emotion. Additionally, with the aim of investigating how the affective features interact with the lexicosemantic properties of words, we performed correlational analyses between norms for familiarity, subjective AoA and SER, and scores for those affective variables which are currently included in the MADs. A distinct pattern of significant correlations with affective features was found for different lexicosemantic variables. These results show that familiarity, subjective AoA and SERs may have independent effects on the processing of emotional words. They also suggest that these psycholinguistic variables should be fully considered when formulating theoretical approaches to the processing of affective language.
Résumé Cet article présente des normes de fréquence subjective pour 660 mots de la langue française recueillies auprès d’adultes jeunes (M = 22,6 ans) et âgés (M = 71,2 ans). La fréquence subjective a été évaluée en utilisant une échelle en 7 points, allant de « jamais rencontré » à « rencontré plusieurs fois par jour ». Les analyses montrent que les estimations sont fidèles pour les 2 groupes d’âge. Les corrélations avec les données issues d’études similaires sont positives et significatives. Par ailleurs, la fréquence subjective corrèle (0,42 à 0,65) avec différents indicateurs de fréquence objective issus de Lexique 3,55 (New et al., 2007). Des analyses de régression indiquent que la fréquence subjective des jeunes adultes est le meilleur prédicteur des performances de décision lexicale d’une population jeune (French Lexicon Project, Ferrand et al., 2010). Enfin, les données indiquent des différences intergénérationnelles dans les estimations pour 24 % des mots. Cette norme, accessible gratuitement ( http://www.labopsycho-u-bordeaux2.fr/psycogni/equipe/cognitive/publis.php?login=robert ), propose un nouvel outil aux chercheurs afin de sélectionner le matériel lexical de langue française utilisé pour étudier les effets liés à l’âge sur le fonctionnement cognitif.
This paper introduces CogNet, a new, large-scale lexical database that provides cognates-words of common origin and meaning-across languages. The database currently contains 3.1 million cognate pairs across 338 languages using 35 writing systems. The paper also describes the automated method by which cognates were computed from publicly available wordnets, with an accuracy evaluated to 94%. Finally, statistics and early insights about the cognate data are presented, hinting at a possible future exploitation of the resource 1 by various fields of lingustics.
Previous research has shown that early-acquired words are produced faster than late-acquired words. Juhasz and colleagues (Juhasz, Lai & Woodcock, Behavior Research Methods, 47 (4), 1004-1019, 2015; Juhasz, The Quarterly Journal of Experimental Psychology, 1-10, 2018) argue that the Age-of-Acquisition (AoA) loci for complex words, specifically compound words, are found at the lexical/semantic level. In the current study, two experiments were conducted to evaluate this claim and investigate the influence of AoA in reading compound words aloud. In Experiment 1, 48 participants completed a word naming task. Using general linear mixed modelling, we found that the age at which the compound word was learned significantly affected the naming latencies beyond the other psycholinguistic properties measured. The second experiment required 48 participants to name the compound word when the two morphemes were presented with a space in-between (combinatorial naming, e.g. air plane). We found that the age at which the compound word was learned, as well as the AoA of the individual morphemes that formed the compound word, significantly influenced combinatorial naming latency. These findings are discussed in relation to theories of the AoA in language processing.
Norms of rated subjective frequency of use and imagery on seven-point scales are reported for 1,916 French nouns. Subjective frequency was defined as the rated frequency of occurrence of words in spoken French, and imagery was defined as the rated case with which a word aroused a mental image. The mean, standard deviation, and percentile rank of the frequency and imagery ratings for each item are presented in the Appendix together with their objective frequency of occurrence in Baudot's (1992) dictionary. Interjudge reliability was assessed by calculating the correlation between the mean ratings of items repeated in the booklet, between the mean ratings obtained from odd-numbered and even-numbered respondents, and by computing the Cronbach alpha statistic for each page of the booklet. These reliability estimates were equal to or greater than.92 for frequency and for imagery, confirming the high level of interjudge consistency. Although the estimates provided by female and male participants were highly correlated (r =.97), the former gave a slightly higher frequency rating to the word sample but a slightly lower imagery rating than the latter did. Moreover, female respondents gave slightly more extreme ratings on the frequency and imagery scales. An analysis of the absolute difference between female and male ratings revealed a discrepancy of one half point or more on 20% of the word sample for frequency and 13% for imagery. On both scales, the mean absolute difference between male and female ratings was larger than that obtained by chance alone. This finding highlights the possibility that some words may not be equally familiar to women and men or may not evoke imagery with the same ease in these groups. Validity estimates for the frequency and imagery ratings were derived from correlations with scale values drawn from other normative studies. These correlation coefficients were equal to or greater than.78 for frequency and.86 for imagery, confirming the high level of consistency between this and other studies. An analysis of the relationship between subjective frequency and imagery ratings indicated that these variables are generally uncorrelated but exceptions occur. In the present study the coefficient of the correlation between subjective frequency and imagery was.24. However, when items with extreme mean frequency were excluded from the calculation, the correlation coefficient dropped to.04 and was no longer significant. Imagery ratings from five independent studies were all positively and significantly correlated with Vikis-Freibergs's (1974) frequency estimates, which were obtained from a free-association task. This finding suggests that word association, as a form of cued recall, may be influenced by several stimulus attributes including prior frequency of association and imagery-evoking value. The pattern of correlation between imagery ratings and text-based frequency estimates is not coherent. It reveals significant correlations only in select cases and no consistent polarity of linear relationship. The main contribution of this research is to provide reliable estimates of subjective frequency and imagery value for a word sample that is larger than those included in previous studies. A close examination of the linear relationship among the various sources of frequency and imagery data underscores the risk of confounding these variables in the selection of lexical stimuli for research.
Cross-Linguistic Norms, Ratings, and Relations for Words and Concepts
We present a collection of concreteness ratings for 35,979 words in Estonian. The data were collected via a web application from 2278 native Estonian speakers. Human ratings of concreteness have not been collected for Estonian beforehand. We compare our results to Aedmaa et al. (2018), who assigned concreteness ratings to 240,000 Estonian words by means of machine learning. We show that while these two datasets show reasonable correlation (R = 0.71), there are considerable differences in the distribution of the ratings, which we discuss in this paper. Furthermore, the results also raise questions about the importance of the type of scale used for collecting ratings. While most other datasets have been compiled based on questionnaires entailing five- or seven-point Likert scales, we used a continuous 0-10 scale. Comparing our rating distribution to those of other studies, we found that it is most similar to the distribution in Lahl et al. (Behavior Research Methods, 41(1), 13-19, 2009), who also used a 0-10 scale. Concreteness ratings for Estonian words are available at OSF.
ANR-funded Nomage project aims at describing the aspectual properties of deverbal nouns taken from a corpus, in an empirical way. It is centered on the development of two resources: a semantically and syntactically annotated corpus of deverbal nouns based on the French Treebank, and an electronic lexicon, providing descriptions of morphological, syntactic and semantic properties of the deverbal nouns found in our corpus. Both resources are presented in this paper, with a focus on the comparison between corpus data and lexicon data.
WordNet is a lexical database for English organized in accordance with current psycholinguistic theories. Lexicalized concepts are organized by semantic relations (synonymy, antonymy, hyponymy, meronymy, etc.) for nouns, verbs, and adjectives.