1396 norm sets
<jats:p>In this paper, we provide an overview of the new GloWbE Corpus — the Corpus of Global Web-based English. GloWbE is based on 1.9 billion words in 1.8 million web pages from 20 different English-speaking countries. Approximately 60 percent of the corpus comes from informal blogs, and the rest from a wide range of other genres and text types. Because of its large size, its architecture and interface, the corpus can be used to examine many types of variation among dialects, which might not be possible with other corpora — including variation in lexis, morphology, (medium- and low-frequency) syntactic constructions, variation in meaning, as well as discourse and its relationship to culture.</jats:p>
<jats:title>Abstract</jats:title><jats:p>Borrowing affixes may be rare compared to lexical borrowing, but it is not random. The current study describes regular patterns of affix borrowing in a database containing 649 borrowed affixes, challenging a number of previous claims about relative borrowability, in particular regarding inflectional categories. It is shown that borrowing affixes of all major nominal and verbal inflectional categories, including case markers and argument indexes, is well attested. Borrowing case markers, for instance, appears to be just as common as borrowing plural markers. By factoring in the “availability” for borrowing (i.e. whether a potential donor language has a relevant affix), it can be shown that nominal categories are far more frequently borrowed than verbal categories. Additionally, it is shown that sets of borrowed affixes often consist of interrelated sets of forms, e.g. forming paradigms, rather than being isolated forms from different morphosyntactic systems, in particular for the more tightly integrated inflectional subsystems. The frequency and systematicity by which inflectional affixes are borrowed calls for a reconsideration of the role of inflection in models of language contact.</jats:p>
<jats:p>RESUMO O presente estudo tem como objetivo descrever os desafios e soluções encontrados na compilação do Corpus de Português Escrito em Periódicos - CoPEP, que contém aproximadamente 40 milhões de palavras, é equilibrado entre as variedades português brasileiro e português europeu em número de palavras e cobre seis grandes áreas de conhecimento. Primeiramente, apresentaremos o contexto de criação do CoPEP, qual seja, a elaboração de um dicionário on-line de português para universitários, para o qual serviu como fonte primária de obtenção de evidências linguísticas. Assim, foram as características desse projeto lexicográfico que informaram os critérios de criação do desenho do CoPEP e as consequentes tomadas de decisão. A seguir, descreveremos a metodologia de aquisição de dados, com foco especial nos desafios enfrentados e nas soluções encontradas. Terminaremos com a descrição da fase final de compilação, na qual aplicamos uma série de procedimentos para obtenção de equilíbrio.</jats:p>
<jats:p>We examined the potential advantage of the lexical databases using subtitles and present SUBTLEX-PT, a new lexical database for 132,710 Portuguese words obtained from a 78 million corpus based on film and television series subtitles, offering word frequency and contextual diversity measures. Additionally we validated SUBTLEX-PT with a lexical decision study involving 1920 Portuguese words (and 1920 nonwords) with different lengths in letters ( M = 6.89, SD = 2.10) and syllables ( M = 2.99, SD = 0.94). Multiple regression analyses on latency and accuracy data were conducted to compare the proportion of variance explained by the Portuguese subtitle word frequency measures with that accounted by the recent written-word frequency database (Procura-PALavras; P-PAL; Soares, Iriarte, et al., 2014). As its international counterparts, SUBTLEX-PT explains approximately 15% more of the variance in the lexical decision performance of young adults than the P-PAL database. Moreover, in line with recent studies, contextual diversity accounted for approximately 2% more of the variance in participants' reading performance than the raw frequency counts obtained from subtitles. SUBTLEX-PT is freely available for research purposes (at http://p-pal.di.uminho.pt/about/databases ).</jats:p>
Feature stability, time and tempo of change, and the role of genealogy versus areality in creating linguistic diversity are important issues in current computational research on linguistic typology. This paper presents a database initiative, DiACL Typology, which aims to provide a resource for addressing these questions with specific of the extended Indo-European language area of Eurasia, the region with the best documented linguistic history. The database is pre-prepared for statistical and phylogenetic analyses and contains both linguistic typological data from languages spanning over four millennia, and linguistic metadata concerning geographic location, time period, and reliability of sources. The typological data has been organized according to a hierarchical model of increasing granularity in order to create datasets that are complete and representative. [ABSTRACT FROM AUTHOR], Copyright of PLoS ONE is the property of Public Library of Science and its content may not be copied o)
The Lesser Sunda Islands in eastern Indonesia cover a longitudinal distance of some 600 kilometres. They are the westernmost place where languages of the Austronesian family come into contact with a family of Papuan languages and constitute an area of high linguistic diversity. Despite its diversity, the Lesser Sundas are little studied and for most of the region, written historical records, as well as archaeological and ethnographic data are lacking. In such circumstances the study of relationships between languages through their lexicon is a unique tool for making inferences about human (pre-)history and tracing population movements. However, the lack of a collective body of lexical data has severely limited our understanding of the history of the languages and peoples in the Lesser Sundas. The LexiRumah database fills this gap by assembling lexicons of Lesser Sunda languages from published and unpublished sources, and making those lexicons available online in a consistent format. T)
Presents a study which aims to investigate SPALEX, a Spanish lexical decision database by focusing on native Spanish speakers at a global scale and with a vast amount of words, to provide a useful tool for researchers exploring the acquisition and processing of this language in native and foreign contexts. SPALEX contains data from a Spanish crowd-sourced lexical decision mega study. The authors collected the data through an online platform from May 12th, 2014 to December 19th, 2017. The majority of the data was acquired during the first month of the experiment, when an advertising campaign was done in order to attract the public’s attention. Participants also had the option of publishing their results via social networks, which led to attract more participants in a snow-ball sampling fashion. Additionally, the database contains information on participants that voluntarily provided information about their gender, age, country of origin, education level, handedness, native language, and best foreign language. In each experimental session, participants responded to 70 words and 30 non-words presented randomly and without repetition. Accuracy in SPALEX is expressed as 1 for correct answers and 0 for incorrect answers. Based on participants’ responses, the authors calculated percentage known, a measure of the percentage of participants that know a particular word. (PsycINFO Database Record (c) 2018 APA, all rights reserved)
We introduce a dataset for studying the evolution of words, constructed from WordNet and the Google Books Ngram Corpus. The dataset tracks the evolution of 4,000 synonym sets (synsets), containing 9,000 English words, from 1800 AD to 2000 AD. We present a supervised learning algorithm that is able to predict the future leader of a synset: the word in the synset that will have the highest frequency. The algorithm uses features based on a word’s length, the characters in the word, and the historical frequencies of the word. It can predict change of leadership (including the identity of the new leader) fifty years in the future, with an F-score considerably above random guessing. Analysis of the learned models provides insight into the causes of change in the leader of a synset. The algorithm confirms observations linguists have made, such as the trend to replace the -ise suffix with -ize, the rivalry between the -ity and -ness suffixes, and the struggle between economy (shorter words ar)
The Moral Foundations Dictionary (MFD) is a useful tool for applying the conceptual framework developed in Moral Foundations Theory and quantifying the moral meanings implicated in the linguistic information people convey. However, the applicability of the MFD is limited because it is available only in English. Translated versions of the MFD are therefore needed to study morality across various cultures, including non-Western cultures. The contribution of this paper is two-fold. We developed the first Japanese version of the MFD (referred to as the J-MFD) using a semi-automated method—this serves as a reference when translating the MFD into other languages. We next tested the validity of the J-MFD by analyzing open-ended written texts about the situations that Japanese participants thought followed and violated the five moral foundations. We found that the J-MFD correctly categorized the Japanese participants’ descriptions into the corresponding moral foundations, and that the Moral F)
Language is one the earliest capacities affected by cognitive change. To monitor that change longitudinally, we have developed a web portal for remote linguistic data acquisition, called Talk2Me, consisting of a variety of tasks. In order to facilitate research in different aspects of language, we provide baselines including the relations between different scoring functions within and across tasks. These data can be used to augment studies that require a normative model; for example, we provide baseline classification results in identifying dementia. These data are released publicly along with a comprehensive open-source package for extracting approximately two thousand lexico-syntactic, acoustic, and semantic features. This package can be applied arbitrarily to studies that include linguistic data. To our knowledge, this is the most comprehensive publicly available software for extracting linguistic features. The software includes scoring functions for different tasks. [ABSTRACT FROM)
<jats:p>Dieser Beitrag soll ein Schlaglicht auf den Status Quo des web-basierten Publizierens in der Linguistik werfen, indem die Entstehung des World Atlas of Language Structures Online (WALS) dokumentiert wird. Parallel zur Veröffentlichung in diesem Blog wurde der Beitrag bei der Tagung Berlin Open '09 eingereicht und akzeptiert. Was ist WALS?</jats:p>
In this study we present some semantic-lexical norms concerning 'fruit' collected from normal subjects. Modelling semantic fluency needs norms for all the given exemplars: at least for Italian language, only a few of them are available. The category 'fruit' is composed of a limited number of exemplars, and some subcategories can be singled out. On a preliminary fluency task, 84 different fruits were produced, and a further normal sample provided ratings for familiarity, prototypicality and their age of acquisition. Moreover the semantic proximity between each pair of the 32 most frequent exemplars was collected and a cluster analysis has been carried out in order to yield an empirical partition of 'fruit' into different subgroups of exemplars. Besides the main models of fluency tasks, some possible advantages offered by these norms in the study of brain-damaged subjects are discussed on the basis of real data obtained from two normal subjects. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
<jats:p>Although retrieval of lexical forms is a prerequisite for language production, research of L2 vocabulary learning has focused much more on meanings and form-meaning mappings than on development of detailed, accessible mental representations of forms. This is particularly true with respect to multi-word items (MWIs). We report an experimental study involving a variety of intra-lexical, usage-based, and interlingual co-determinants of L2 vocabulary learnability pertaining to MWIs. Each learner (N = 60) encountered a randomly allocated set of 26 two-word MWIs (Nsets = 4) semi-randomly drawn from a larger pool of MWIs. Learners were asked to remember either the 13 MWIs showing the form variable assonance (e.g., <jats:italic>change shape</jats:italic>) or the 13 nonassonant control MWIs (e.g., <jats:italic>sound good</jats:italic>). Posttests of form recall revealed a large, durable effect of the focusing task in combination with forewarning of testing. Except when MWI concreteness (a semantic variable) was high, assonance had a positive effect on retrievability in recall tests given after delays of 15 minutes and one week. There was a consistent effect of the semi-semantic variable Mutual Information. Even in the context of a strong focus on forms, form variables are not the only variables that matter.</jats:p>
Background: The Object and Action Naming Battery (OANB) was developed by Druks and Masterson in 2000 in response to the lack of materials for investigating the difference between the availability of nouns and verbs. This battery has been extensively used in psycholinguistic and aphasia research. The battery has also proved to be a useful tool in clinical practice by speech and language therapists. Aims: Till date, there are no published aphasia assessment tools specifically developed for the use of Saudi Arabic speakers. Therefore, the present study aimed to adapt the OANB for the use of Saudi Arabic speakers. This paper describes the adaptation process. Methods & Procedures: Name agreement data for the items in the OANB was collected from 30 non-brain-damaged Saudi Arabic-speaking adults. This was followed by collecting values for the psycholinguistic variables available in the original battery, which are spoken-word frequency, imageability, age of acquisition, and visual complexity. Outcomes & Results: The Saudi Arabic version of the OANB consists of 50 object and 50 action pictures with high level of name agreement (100% for object pictures, and at least 93% for action pictures), along with the normative data for the variables of spoken-word frequency, imageability, age of acquisition, and visual complexity of the verbal labels for the object and action pictures included in the Saudi Arabic version of the battery. Conclusions: This battery makes a significant contribution to aphasia resources available in Saudi Arabia as it can be used in clinical settings at the assessment stage and for therapeutic purposes for individuals with aphasia. The battery can also be used in aphasia and psycholinguistic research with Arabic speakers.
INTRODUCTION: The ability to name pictures has been investigated widely in healthy people and clinical populations. The Object and Action Naming Battery (OANB) is widely used for psycholinguistic research, aphasia research, and clinical practice. Normative databases for pictorial stimuli have been conducted in language processing studies to control for various psycholinguistic variables known to affect the availability of picture names. The present study provides Moroccan Arabic norms for name agreement, familiarity, imageability, visual complexity, and age of acquisition for 100 line drawings of actions and 162 line drawings of objects taken from Druks and Masterson. METHODS AND PROCEDURES: 160 healthy Moroccan Arabic-speaking individuals participated in this study. Name agreement values for the OANB items were collected from forty subjects, followed by collecting data for the psycholinguistic variables: spoken-word frequency, imageability, visual complexity, and age of acquisition from 120 participants. RESULTS: The Moroccan Arabic OANB (MA-OANB) comprises 70 objects and 60 action pictures. 77% of the nouns and 68% of the verbs obtained 100% target responses. A minimum of 93 percent name agreement was reached for the remaining items. Norms were also collected for the following psycholinguistic variables: spoken-word frequency, imageability, age of acquisition, and visual complexity. CONCLUSION: The stimuli can be used for various psycholinguistic investigations and also for assessment and therapeutic purposes in Morocco.
The field of false memories has been widely studied in cognitive psychology through the DRM paradigm (Deese, 1959;Roediger & McDermott, 1995), an experimental task used to induce false memories from materials that are conceptually and semantically related. This paradigm involves presenting lists of words that are semantically related to a non-presented critical word, with a subsequent memory test showing high levels of false recall and false recognition of that non-studied critical word. For example, after studying a list containing words, such as "butter", "food" and "sandwich", it is likely that in a subsequent free recall or recognition test, the word "bread" will be mistakenly identified as studied.False memory is generated, in part, by the relationship between the list words and the critical word (Gallo, 2010;Roediger & Gallo, 2016). This phenomenon has been explained by two theories: the fuzzy-trace theory (FTT) (Brainerd & Reyna, 1998), and the activationmonitoring framework (AMF) (Roediger et al., 2001). Both theories agree that, understanding the production of false memories requires considering two complementary processes: an error inflation process, identified as gist encoding in FTT and as activation in AMF; and an error editing process, identified as monitoring in AMF and as recollection rejection in FTT (Brainerd & Reyna, 2002).According to FTT (Brainerd & Reyna, 1998), "gist" refers to the general theme extracted from studied material. When a list of words related to a non-presented critical word is studied, both literal and semantic information are encoded. In a subsequent memory test, these literal and semantic memory traces operate simultaneously, providing information about the items. The retrieval of semantic information may lead to considering the critical word as having been studied due to its similarity to the presented words. However, the retrieval of literal information about the studied associates can counteract this effect by providing evidence that the critical word was not studied, a process known as recollection rejection (Brainerd & Reyna, 2002).According to the AMF (Roediger et al., 2001), two processes work together to produce false memory: activation and monitoring. Studying a list of words can trigger activation that spreads through the lexical-semantic system, creating implicit associations between interconnected words. This activation is moderated by a subsequent monitoring process that helps distinguishing between correct recall (studied words) and false memories (non-studied words).Even when considering different perspectives, both FTT and AMF agree that presenting a list of associates words activates a non-presented critical word, leading to error inflation. If the error inflation process is not accompanied by its corresponding error editing process, or if this process fails, false recall or false recognition may occur (Arndt & Gould, 2006). Therefore, studying the strategies used to avoid false memories is crucial for understanding the mechanisms underlying their formation.Theme identifiability of a list is one of the essential factors involved in this error editing process (Carneiro et al., 2009). In their normative study in Portuguese, they provided theme identifiability norms for 40 DRM associative lists selected from Albuquerque's (2005) study. Participants were presented with the lists and asked to generate a word that best described the general theme of each list. To study the effect of this factor on the production of false memories, they selected those lists with the highest and lowest levels of theme identifiability. The results showed that lists with high identifiability of the critical word as the theme produced lower levels of false recall and false recognition compared to lists where the critical word was not as easily identifiable. According to Carneiro et al. (2012), this outcome was attributed to an error editing strategy called "Identify to reject," which involves several stages: detecting that all the words in a list are related to a common theme, identifying the word that best describes the theme but is not present in the list, keeping it in mind to avoid recalling it in the future, and consequently, reducing false memories. This pattern has been observed in other studies using lists with an associative structure (Beato et al., 2023;Carneiro & Fernandez, 2013;Carneiro et al., 2012). Additionally, there are other normative studies that provide theme identifiability indices for associative lists in Spanish (Beato & Cadavid, 2016) and in English (Neuschatz et al., 2003).Most studies on false memories and error editing mechanisms using the DRM paradigm employ lists with an associative structure. However, ad hoc categorical relationships are less explored in the literature. Ad hoc categories are spontaneously constructed to achieve a specific goal in a given context, and their elements can come from different taxonomic categories (e.g., "Things that can fall on your head") (Barsalou, 1983). Both common and ad hoc categories can lead to similar memory distortions. While false memories are typically more pronounced for common categories, they are still robust for ad hoc categories (Soro et al., 2017).The mechanisms for avoiding memory distortions in such lists are not well understood. Ad hoc categories provide a valuable tool for studying situated concept representations, which are characterized by their flexibility and dynamism (Barsalou, 2005). Their use allows researchers to explore how individuals organize and retrieve information when categories are not predefined but emerge from context. This is particularly relevant for understanding how memory adapts to new information and situations, adding depth to theoretical debates on flexible concept representation.In the study with associative lists by Carneiro et al. (2009), the percentages of identification for critical words ranged from 1% to 77%. However, Soro et al. (2017) indicate that, unlike associative lists, in ad hoc categorical structured lists, theme identification typically refers to identifying the category label rather than the critical word itself. They used two criteria: exact identifiability, where participants identified the original theme of the lists (e.g., "Materials that cover the ground" for the category "Things that can be walked upon"), and comprehensive identifiability, where participants identified a theme that could include the critical word (e.g., a label that includes the critical word "grass" for "Things that can be walked upon"). Both criteria are important for describing our findings.To our knowledge, no previous studies have addressed false memories or theme identification with ad hoc categories in Spanish. Therefore, the aim of this research was to obtain theme identifiability indices for 70 lists that maintained ad hoc categorical relationships with a non-presented critical word using the DRM paradigm. These lists were created based on a normative study of ad hoc categories conducted in Spanish, which, to our knowledge, is the first of its kind in this language. Additionally, this research aims to lay the groundwork for studying the underlying mechanisms of error editing processes in lists with ad hoc categorical relationships in Spanish.In future research, these data may help in understanding the role of theme identifiability in the formation of false memories. Moreover, having these indices will enable more accurate predictions and better control over the experimental materials.A total of 188 students from the Psychology degree program at the University of La Laguna participated. All participants were native Spanish speakers (146 women, 42 men; mean age= 20.14, SD = 2.94).The material consisted of 70 ad hoc categorical lists, each containing 10 words. Both the critical words and their corresponding associates were extracted from a normative study conducted to collect data on ad hoc categories in Spanish (Alonso et al., in preparation;Benítez, Alonso, Fernandez, & Díez, 2022). Ad hoc categories were selected from various normative studies in English and Portuguese and then translated into Spanish (Barsalou, 1982(Barsalou,, 1983(Barsalou,, 1985;;Hough & Pierce, 1989;Soro & Ferreira, 2017;Vallée-Tourangeau et al., 1998;van Overschelde et al., 2004). The general procedure used in the Spanish normative study was similar to that used by Battig & Montague (1969), with the exception that in the present study both the presentation of the material and the collection for responses were done by computer (see van Overschelde et al., 2004). The participants were instructed to generate as many exemplars as possible for each category within one minute. This study, currently in preparation, will provide indices of frequency, rank, and lexical availability for the exemplars of each category.The lists were constructed based on the frequencies obtained from the normative study. Critical words were selected as those with the highest frequency within each category, while their associates were the next most frequent words. Care was taken to ensure that the critical words did not appear in more than one list and that no associate was repeated across lists. The selected critical words were primarily nouns (with only 4 being verbs and 1 adjective), ranging from 1 to 5 syllables, with a mean frequency of occurrence in Spanish of 60.86 per million (Alonso et al., 2011).The 70 lists were divided into 5 blocks: 4 blocks containing 15 lists each and 1 block with 10 lists. For the theme identification test, a booklet was prepared with several pages. The first page collected participant information (name, age, gender, and degree). The second page included practice examples to familiarize participants with the task. The remaining pages were dedicated to the experimental lists. Each page displayed the list number and had three blank spaces for participants to write the word or words they believed identified the theme of the list (up to three), along with a Likert scale to indicate their confidence in whether the word given in the first position represented the list's theme. Procedure Data collection took place in November 2023 during group sessions, each consisting of approximately 35 participants and lasting around 30 minutes. Each group studied 15 lists, except for one group that studied only 10 lists. The order of list presentation within each group was randomly determined.Participants began by completing the demographic information on the first page of the booklet and then received instructions similar to those used in previous studies (Carneiro et al., 2009;Neuschatz et al., 2003). They were shown a series of lists in a PowerPoint presentation, with one word displayed every 2 seconds. Before the start of the experimental session, and to ensure participants fully understood the task, two practice trials were conducted. These trials involved the same task as the main session: following the presentation of each list, participants had 50 seconds to generate up to three words that they believed best described the theme of the list. They also provided a confidence judgment on how certain they were that the first word given represented the list's theme, using a Likert scale ranging from 1 ("not very confident") to 5 ("very confident").Before each list, a message appeared on the screen indicating the list number to be presented. The process of presenting the list and identifying the theme was repeated until the experimental session was complete.The spreadsheet file accompanying this report (Theme_Identifiability_Ad_hoc_Spanish.xls) consists of five sheets. The first sheet, "Theme identifiability", contains the raw data for each participant. It includes the 70 critical words and their respective lists of 10 ad hoc associates. The first column shows the participant number, the second column identifies the list, the third column specifies the type of relationship of the lists (ad hoc), and the fourth column indicates the type of word (studied vs. critical). The fifth column contains the words themselves, while the sixth column provides the English translations of the critical and studied words. Adjacent columns include all responses provided by participants in the first, second, and third positions, as well as the confidence judgements related to the first response. Additionally, intrusions are noted-i.e., words from a study list that a participant mistakenly identified as the theme of that list, and therefore are not considered valid responses.The second sheet, "First word", summarizes the total count of responses given in the first position for each list. It includes all the words generated by participants as the theme in the first position, associated with their respective critical word and list number. Additionally, it provides the English translation of the critical word, the total number of participants who responded to each list, the number and percentage of participants who identified a word as the theme, and the mean confidence judgement for each first-word response. Intrusions are also noted, including quantity and incorrectly identified words.The third and fourth sheets, "Second word" and "Third word", respectively, are dedicated to responses given in the second and third positions. The layout is identical to the previous sheets, but it does not include the column for mean confidence judgement. The fifth sheet, "Summary," presents the final summary, showing the most frequently identified theme for the first, second, and third positions for each list.Responses recorded in the database were maintained in their original format, with corrections made only for spelling errors. Singular/plural and masculine/feminine forms of words were counted separately.Table 1, available as supplementary material, shows the 70 critical words with their corresponding list identifier and English translation, the number of participants who responded to each list, the number of different themes given as the first response, the percentage of participants who indicated the critical word (comprehensive identification) as the theme of the list in the first position, and the mean confidence rating for the critical word as the first response. Additionally, it provides the percentages of participants who identified the ad hoc category label (exact identification) in the first position, as well as the mean confidence rating for these first-position responses.Theme identifiability is a crucial factor in the error editing process and thus influences the formation of false memories (Carneiro et al., 2009;Soro et al., 2017). However, the study of false memories in ad hoc categorical lists and the factors contributing to their formation remain unexplored in Spanish. Therefore, this study aimed to provide theme identifiability indices for 70 ad hoc lists within the framework of the DRM paradigm in Spanish.When examining identifiability levels using the same approach as in studies with associative lists (Carneiro et al., 2009), where the critical word is considered as the theme, the levels of identifiability are relatively low. Regarding comprehensive identifiability, the critical word was identified as a theme in the first position in only 10 out of the 70 lists, with identification percentages ranging from 2.1% to 17.2%. On average, the critical word was identified as the theme 0.98% (SD=3.11) of the time in the first position across all 70 lists. When considering only where the critical word was identified, this average increased to 6.88% (SD=5.39), with a mean confidence rating of 3.57 (SD=1.29). Given the ad hoc category exemplars can come from different categories, their membership is not immediately apparent without context, making the activation of the critical word challenging.In contrast, exact identifiability showed higher levels of identification. The theme of 59 lists (ranging from 2.1% to 87.5%) was identified in the first position, demonstrating a significant increase in identification compared to when only the critical word was considered. Participants often generated words that, while not precisely matching the category labels, were related to them, suggesting some thematic processing even if not explicitly expressed. On average, the theme was identified 22.22% (SD=23.75) of the time in the first position across all 70 lists. When considering only the lists where the theme was identified, this average rose to 26.36% (SD=23.67), with a mean confidence rating of 3.98 (SD=0.70) for the first response.According to Carneiro et al. (2009), in associative lists, the level of identification of the critical word as the theme is inversely related to the occurrence of false memories. Identifying the critical word as the theme triggers an error editing process that mitigates false memories in subsequent memory tests. However, this assumption may not hold true for categorical lists, particularly ad hoc categorical lists, where theme identifiability might not exert the same effect on the error editing process. Soro et al. (2017) suggested that for false memories to occur with ad hoc categories, a positive relationship with theme identifiability might be necessary due to the inherent variability among exemplars. Their study found no correlation between false recognition and theme identifiability. Instead, false recognition was influenced more by individual factors such as experience or creative thinking. Some participants generated more associations between exemplars, leading to an increased likelihood of errors in a subsequent recognition test.Our results suggest that context plays a fundamental role in theme identification within ad hoc categories, particularly concerning comprehensive identifiability. Therefore, considering the findings of Soro et al. (2017), the participant's ability to integrate exemplars into an appropriate context may significantly influence the activation of critical words. This aligns with the notion that memory is a dynamic process shaped by contextual factors (Barsalou, 2005).The present study provides valuable data on theme identifiability for a wide range of ad hoc categorical lists, underscoring the need for a different approach when studying error editing processes in this context. Our findings reveal a crucial aspect of ad hoc categories: while traditional associative lists tend to exhibit reduced false memories due to clearer theme identifiability, the inherent variability and contextual emergence of ad hoc categories introduce complexities in memory retrieval that merit further investigation. This area represents a promising opportunity for advancing our understanding of false memories.
In this norming study, for 2100 Serbian nouns, we collected ratings on familiarity, concreteness, imageability, age of acquisition, context availability, emotional valence, arousal, the possibility of experiencing a concept on each of the five sensory modalities (visual, auditory, olfactory, gustatory, tactile), and the extent of the actual experience for the same concept. Based on the sensory ratings, the integrative perceptual richness measures were derived: maximal perceptual strength, modality exclusivity, number of modalities, the sum of ratings, Euclidian vector length, and Minkowski 3 distance. The principal component analysis revealed different factor structure for the perceptual strength measures based on the possible and the real experience. All modalities except the auditory grouped into one component for a possible experience. On the other hand, PCA analysis for the real experience ratings showed that the gustatory and olfactory modality migrated into a separate dimension, suggesting that concepts are less multimodal and described mainly by visual/tactile olfactory/gustatory experience. Finally, we conducted a lexical decision over the entire data set of words. Gustatory, olfactory and tactile strength significantly accelerated word processing when estimates were based on the possible experience. When estimates were grounded on real experience, auditory strength inhibited processing additionally. Analysis of the perceptual richness measures showed that the modality exclusivity was the only measure with the consistent inhibitory effect in all analyses. Because the high values of modality exclusivity indicated the unimodality's tendency, our results showed that words perceived with only one modality took more time to process.
Aims and objectives: English has become the dominant donor language for many languages, including Croatian. Perception of English loanwords has mainly been investigated through corpus-based studies or attitude questionnaires. At the same time, normative data for unadapted English loanwords are still mainly unavailable. This study aims to fill that gap by collecting affective and lexico-semantic norms for unadapted English loanwords in Croatian. Methodology: Valence, arousal, familiarity, and concreteness ratings for unadapted English loanwords and three types of Croatian equivalents were collected from 565 participants. Data and analysis: Affective and lexico-semantic norms for each word on the four variables are available in the database. In addition, the relationship between different variables was examined. Finally, the differences between English loanwords and three types of Croatian equivalents (in-context, out-of-context, and adapted forms) are reported. Findings: Valence ratings for unadapted English loanwords differed from out-of-context equivalents and adapted forms. Unadapted English loanwords were rated as more arousing than Croatian equivalents. Finally, unadapted English loanwords were less familiar and less concrete than in-context and out-of-context equivalents. The findings suggest that Croatian speakers perceive unadapted English loanwords differently on affective and lexico-semantic levels compared with Croatian equivalents. Originality: This is the first study to provide affective and lexical norms for 391 most frequent unadapted English loanwords in Croatian. Implications: The reported normative data will contribute to the existing knowledge about the processing of English loanwords by enabling experimental research on this topic.
Abstract: Subjective ratings of dimensions of lexical meaning have long been used in experimental psychology and psycholinguistics―for example, in experimental studies of memory, lexical processing, and brain function. Three such dimensions of lexical meaning are concreteness (vs abstractness), emotional valence (degree of pleasantness), and arousal (degree of excitement). Ratings have typically been obtained by presenting lexical items to multiple respondents who rate each item on a Likert scale, after which the ratings for each item are averaged. For some dimensions of meaning and for some languages, ratings of many thousands of single words are freely available; but even for English there are as yet no remotely similar-size collections of ratings for multiword expressions (MWEs), such as collocations and spaced compound nouns. Researchers, including researchers of L2 vocabulary acquisition, may therefore wonder how well a MWE’s level of concreteness, valence or arousal can be estimated from the ratings of its constituent words. This article reports a study which addressed that question, concluding that MWE ratings derived from constituent word ratings must be used with caution. The study has doubled the e)xisting small stock of English MWEs rated for arousal and valence. The data: The spreadsheet relates to seven correlational studies of which five concern concreteness and one each concern valence and arousal. The focal data in each substudy consist of a column of subjective ratings of (a) MWEs as wholes and (b) mean ratings, i.e., (rating of word 1 + rating of word 2) / 2.The concreteness ratings mostly come from Brysbaert, Warriner, and Kuperman (2014) although for study two some of the word ratings come from the MRC Psycholinguistic Database (Coltheart, 1981; Wilson, 1988). The great majority of the MWE ratings of valence and arousal were collected by me through Amazon Mechanical Turk; some stem from Warriner, Kuperman, and Brysbaert (2013). All the word ratings stem from Warriner et al. An additional crucial variable in one study (Study 6, Valence) is the valence rating of the most-valenced constituent word--i.e., the word whole rating departs most in either direction from neutral, which is 5 on the 9-point Likert scale of the valence and arousal ratings. The concreteness ratings are on a 5-point scale. Among the data for valence and arousal are ratings for items for which WKB and AMT ratings are available. Correlations between these WKB and AMT items furnish some degree of validation. The data are described more fully in a soon-to-be-submitted article entitled, 'Measuring perceptual and emotive dimensions of multi-word expressions: Are constituent word ratings enough?' <br>Abbreviations: C-word = Constituent word; BWK = Brysbaert et al.; WKB = Warriner et al.; AMT = Amazon Mechanical Turk; MRC = The MRC Psycholinguistic Database<br>ReferencesBrysbaert, M, Warriner, A, and Kuperman, V (2014) Concreteness ratings for 40,000 generally known English word lemmas. Behavior Research Methods 46: 904–11. Coltheart, M (1981) The MRC Psycholinguistic Database, Quarterly Journal of Experimental Psychology 33A: 497–505. Warriner A, Kuperman, V, and Brysbaert, M (2013) Norms of valence, arousal, and dominance for 13,915 English lemmas. Behavior Research Methods 45: 1191–207. List retrieved from: http://crr.ugent.be/archives/1003Wilson, M. (1988). The MRC Psycholinguistic Database: Machine readable dictionary, Version 2, Behavioural Research Methods, Instruments and Computers 20: 6-11. Retrieved from: http://websites.psychology.uwa.edu.au/school/MRCDatabase/uwa_mrc.htm<br><br>