1358 norm sets
Studies on object and word naming have shown that the age at which words are acquired is an important factor in processing times. Research on the issue in Dutch has been hampered by the fact that only teacher ratings were available about which words should be known by 6-year-olds. As a supplement to these teacher ratings, we conducted a large-scale study in which 558 students rated the age-of-acquisition of 2816 four- and five-letter nouns. Reliability of the ratings is high, and correlations with word frequency and word imageability are in the same order as those reported for English.
Based on a sample of 145 Flemish first year psychology students at the University of Leuven (Belgium), affective and subjective familiarity norms were obtained for 740 Dutch words. One group of students (N = 64) rated 370 nouns, and a second group (N = 81) rated 370 personality-trait words on seven-point visual analogue scales, both for the positive-negative, and familiar-unfamiliar dimensions. Test-retest and inter-rater reliability coefficients were very high for both wordsets and response-types. The mean ratings and their standard deviations are presented in the Appendix. Gender differences for specific words are tabulated, and the observed association between the affective and the familiarity ratings is discussed.
This study presents subjective ratings for 3,022 Croatian words, which were evaluated on two affective dimensions (valence and arousal) and one lexico-semantic variable (concreteness). A sample of 933 Croatian native speakers rated the words online. Ratings showed high reliabilities for all three variables, as well as significant correlations with ratings from databases available in Spanish and English. A quadratic relation between valence and arousal was observed, with a tendency for arousal to increase for negative and positive words, and neutral words having the lowest arousal ratings. In addition, significant correlations were found between affective dimensions and word concreteness, suggesting that abstract words have a tendency to be more arousing and emotional than concrete words. The present database will allow experimental research in Croatian, a language with a considerable lack of psycholinguistic norms, by providing researchers with a useful tool in the investigation of the relationship between language and emotion for the South-Slavic group of languages.
Normative studies are common in cognitive psychology because they allow us to estimate with more precision the attributes of the stimuli used in empirical studies. The studies reported here had four aims. The first three aims were to obtain estimates for (a) familiarity, concreteness, valence, and arousal for a single set of words in Brazilian Portuguese; (b) wordlikeness (similarity to Portuguese) of a set of foreign words (Swahili); and (c) recall accuracy of Swahili–Portuguese word pairs in a multitrial learning task. The fourth aim was to investigate if any of the assessed measures predicts recall accuracy. One-hundred twenty-eight participants took part in one of the three studies. In Studies 1a and 1b, participants judged 80 Portuguese words for familiarity, concreteness, valence, and arousal and 80 corresponding Swahili words for wordlikeness; in Study 2, participants carried out three study–test cycles of a set of Swahili–Portuguese word pairs. Overall, word-attribute estimates were reliable ( rs = .94–.98) and participants’ responses had high internal consistency (Cronbach’s α = .84–.98). Moreover, the relative difficulty of word pairs was retained across trials ( rs = .65–.88). Although different variables correlated with recall accuracy at different time points, multiple regressions indicate that none of the word-attribute variables predicted recall accuracy across trials. These norms may prove fruitful not only for Brazilian human memory researchers but also for international research teams, as it will enable the development of more controlled cross-cultural studies in this field.
Researchers have only recently started to take advantage of the developments in technology and communication for sharing data and documents. However, the exchange of experimental material has not taken advantage of this progress yet. In order to facilitate access to experimental material, the Bank of Standardized Stimuli (BOSS) project was created as a free standardized set of visual stimuli accessible to all researchers, through a normative database. The BOSS is currently the largest existing photo bank providing norms for more than 15 dimensions (e.g. familiarity, visual complexity, manipulability, etc.), making the BOSS an extremely useful research tool and a mean to homogenize scientific data worldwide. The first phase of the BOSS was completed in 2010, and contained 538 normative photos. The second phase of the BOSS project presented in this article, builds on the previous phase by adding 930 new normative photo stimuli. New categories of concepts were introduced, including animals, building infrastructures, body parts, and vehicles and the number of photos in other categories was increased. All new photos of the BOSS were normalized relative to their name, familiarity, visual complexity, object agreement, viewpoint agreement, and manipulability. The availability of these norms is a precious asset that should be considered for characterizing the stimuli as a function of the requirements of research and for controlling for potential confounding effects.
Despite the flourishing research on the relationships between affect and language, the characteristics of pain-related words, a specific type of negative words, have never been systematically investigated from a psycholinguistic and emotional perspective, despite their psychological relevance. This study offers psycholinguistic, affective, and pain-related norms for words expressing physical and social pain. This may provide a useful tool for the selection of stimulus materials in future studies on negative emotions and/or pain. We explored the relationships between psycholinguistic, affective, and pain-related properties of 512 Italian words (nouns, adjectives, and verbs) conveying physical and social pain by asking 1020 Italian participants to provide ratings of Familiarity, Age of Acquisition, Imageability, Concreteness, Context Availability, Valence, Arousal, Pain-Relatedness, Intensity, and Unpleasantness. We also collected data concerning Length, Written Frequency (Subtlex-IT), N-Size, Orthographic Levenshtein Distance 20, Neighbor Mean Frequency, and Neighbor Maximum Frequency of each word. Interestingly, the words expressing social pain were rated as more negative, arousing, pain-related, and conveying more intense and unpleasant experiences than the words conveying physical pain.
Background: The embodied cognition approach, as applied to concrete knowledge, is centred on the role of the perceptual and motor aspects of experience. To extend the embodied framework to abstract knowledge, some studies have suggested that further dimensions, such as affective or social experiences, are relevant for the semantic representations of abstract concepts. The objective of this study is to develop a measure that can quantitatively capture the multidimensional nature of abstract concepts. Methods: We used dimension-rating methods, known to be suitable, to account for the semantic representations of abstract concepts, to develop a new database of 964 Italian words, rated by 542 participants. Besides classical psycholinguistic variables (i.e., concreteness, imageability, familiarity, age of acquisition, semantic diversity) and affective norms (i.e., valence, arousal), we collected ratings on selected dimensions characterizing the semantic representations of abstract concepts, i.e., introspective, mental state, quantitative, spatial, social, moral, theoretical, and economic dimensions. The measure of exclusivity was incorporated to quantify the number of dimensions, and the respective relevance, for each concept. Concepts with a high value of exclusivity rely on only one/a few dimension/s with high value on the respective rating scale. Results: A multidimensional representation characterized most abstract concepts, with two robust major clusters. The first was characterized by dense intersections among introspective, mental state, social, and moral dimensions; the second, less interconnected, cluster revolved around quantitative, spatial, theoretical, and economic dimensions. Quantitative, theoretical, and economic concepts obtained higher exclusivity values. Conclusions: The present study contributes to the investigation of the semantic organization of abstract words and supports a controlled selection and definition of stimuli for clinical and research settings.
The organization of abstract concepts reflects different dimensions, grounded in the brain regions coding for the corresponding experience. Normative measures of linguistic stimuli offer noteworthy insights into the organization of conceptual knowledge, but studies differ in the dimensions and classes of concepts considered. Additionally, most of the available information has been collected in English, without considering possible linguistic and cultural differences. Here, we aimed to create a comprehensive Turkish database for abstract concepts (TACO), including rarely investigated classes such as political concepts. We included 503 words-78 concrete (fruits, animals, tools) and 425 abstract (emotions, social, mental states, theoretical, quantity, space, political)-rated by 134 Turkish speakers for familiarity, imageability, age of acquisition, valence, arousal, quantity, space, theoretical, social, mental state, and political dimensions. We calculated dominance and exclusivity, indicating the dimension receiving the highest mean score for each word, and the position of the word along the unidimensional–multidimensional continuum, respectively. A principal component analysis (PCA) was conducted on the semantic dimensions. The results showed that mental state was the dominant dimension for most concepts. Moderate to low levels of exclusivity indicated that the concepts were multidimensional. PCA revealed three components: Component 1 captured the juxtaposition between social/mental state and magnitude polarities, Component 2 highlighted affective components, and Component 3 grouped together political and theoretical dimensions. The introduction of political concepts provided insights into the multidimensional nature of this unexplored class, closely intertwined with the theoretical dimension. TACO constitutes the first comprehensive Turkish database covering several abstract dimensions, paving the way for cross-linguistic and cross-cultural studies of semantic representations.
Project description This project hosts TUNorms, a set of psycholinguistic word norms for Thai developed to support research on lexical–semantic processing and embodied cognition. The dataset provides normative ratings for imageability, body–object interaction (BOI), and subjective frequency for 627 mono- and multi-syllabic Thai words. The norms were collected from Thai university students using standardised rating procedures. Reliability was assessed through internal consistency and cross-linguistic comparisons with existing norms for overlapping items, and the measures were further validated in a semantic categorisation task demonstrating independent effects of imageability and BOI on response latencies after controlling for established lexical variables. The repository contains the aggregated normative data and the full rating instructions (Thai and English). Raw participant-level data are not included, as this is a completed normative study and the shared materials are intended to support controlled stimulus selection, replication, and secondary analyses. This dataset accompanies a journal manuscript currently under submission. Upon acceptance, the final citation will be added to this record. The norms are released under a Creative Commons Attribution 4.0 International (CC BY 4.0) licence to facilitate reuse.
There are currently stimuli with published norms available to study several psychological aspects of language and visual cognitions. Norms represent valuable information that can be used as experimental variables or systematically controlled to limit their potential influence on another experimental manipulation. The present work proposes 480 photo stimuli that have been normalized for name, category, familiarity, visual complexity, object agreement, viewpoint agreement, and manipulability. Stimuli are also available in grayscale, blurred, scrambled, and line-drawn version. This set of objects, the Bank Of Standardized Stimuli (BOSS), was created specifically to meet the needs of scientists in cognition, vision and psycholinguistics who work with photo stimuli.
Humans have a remarkable fidelity for visual long-term memory, and yet the composition of these memories is a longstanding debate in cognitive psychology. While much of the work on long-term memory has focused on processes associated with successful encoding and retrieval, more recent work on visual object recognition has developed a focus on the memorability of specific visual stimuli. Such work is engendering a view of object representation as a hierarchical movement from low-level visual representations to higher level categorical organization of conceptual representations. However, studies on object recognition often fail to account for how these high- and low-level features interact to promote distinct forms of memory. Here, we use both visual and semantic factors to investigate their relative contributions to two different forms of memory of everyday objects. We first collected normative visual and semantic feature information on 1,000 object images. We then conducted a memory study where we presented these same images during encoding (picture target) on Day 1, and then either a Lexical (lexical cue) or Visual (picture cue) memory test on Day 2. Our findings indicate that: (1) higher level visual factors (via DNNs) and semantic factors (via feature-based statistics) make independent contributions to object memory, (2) semantic information contributes to both true and false memory performance, and (3) factors that predict object memory depend on the type of memory being tested. These findings help to provide a more complete picture of what factors influence object memorability. These data are available online upon publication as a public resource.
Auditory pseudowords are widely used in psycholinguistics and cognitive neuroscience, but their construction requires control of sublexical familiarity and careful characterization of how acoustic cue manipulations may shift perceived lexical plausibility. Here we introduce the Minho Pseudoword Wordlikeness Ratings (MPWR), the first normative dataset of wordlikeness judgments for European Portuguese (EP) auditory trisyllabic CV pseudowords, and evaluate whether adding a localized F0-based prominence cue modulates wordlikeness beyond distributional familiarity. One hundred and twenty pseudowords were assembled from naturally produced syllables drawn from the Minho Spoken Syllable Pool (MSSP) and recorded under uniform conditions. Each item was implemented in three token types with constant segmental content: a flat baseline and two F0-enhanced versions (+15%) targeting either the penultimate or final syllable. Native EP listeners (N = 101) provided wordlikeness ratings on a 7-point scale. MSSP-derived indices quantified pseudoword syllable familiarity (SWIAll, SWIN3) and stress-position propensity for the targeted syllable (SPPmarked). Ratings were intentionally low overall yet showed substantial item-to-item variability. F0 enhancement produced a small but reliable decrease in wordlikeness relative to flat tokens, with no reliable difference between penultimate and final targeting positions. SWIAll robustly predicted ratings, whereas SPPmarked added little explanatory value. MPWR provides a practical EP resource for selecting and matching auditory pseudowords using normative wordlikeness ratings and transparent corpus-based descriptors.
Normed linguistic stimuli are fundamental in psycholinguistics because they capture lexical and semantic properties that influence comprehension. However, generating these norms at scale is challenging, often leading researchers to rely on ad hoc norms collected from small samples, which can introduce inconsistencies and limit cross-study comparisons. In the present study, we investigated how large language models (LLMs) can support psycholinguistic research by prompting eight current LLMs to norm 300 English two-word metaphor combinations, such as sharp mind. We selected the dimensions of familiarity, aptness, concreteness, metaphoricity, and constituency, as these tap distinct cognitive processes and may provide insight into which aspects LLMs capture accurately and which they do not. We varied stimulus presentation (in context vs. in isolation) and response format (categorical vs. numerical) to examine which manipulation yields norms most closely aligned with human ratings. We then assessed the reliability and validity of model responses and used them to replicate existing analyses of metaphor comprehension. Overall, LLM-generated norms aligned best with familiarity and metaphoricity, which rely on word co-occurrence. In contrast, aptness, concreteness, and constituency—which require reasoning about the relationship between the topic (e.g., mind) and the vehicle (e.g., sharp)—proved more challenging for LLMs.
Normed linguistic stimuli are fundamental in psycholinguistics because they capture lexical and semantic properties that influence comprehension. However, generating these norms at scale is challenging, often leading researchers to rely on ad hoc norms collected from small samples, which can introduce inconsistencies and limit cross-study comparisons. In the present study, we investigated how large language models (LLMs) can support psycholinguistic research by prompting eight current LLMs to norm 300 English two-word metaphor combinations, such as sharp mind. We selected the dimensions of familiarity, aptness, concreteness, metaphoricity, and constituency, as these tap distinct cognitive processes and may provide insight into which aspects LLMs capture accurately and which they do not. We varied stimulus presentation (in context vs. in isolation) and response format (categorical vs. numerical) to examine which manipulation yields norms most closely aligned with human ratings. We then assessed the reliability and validity of model responses and used them to replicate existing analyses of metaphor comprehension. Overall, LLM-generated norms aligned best with familiarity and metaphoricity, which rely on word co-occurrence. In contrast, aptness, concreteness, and constituency—which require reasoning about the relationship between the topic (e.g., mind ) and the vehicle (e.g., sharp )—proved more challenging for LLMs.
Emotion lexicons are useful in research across various disciplines, but the availability of such resources remains limited for most languages. While existing emotion lexicons typically comprise words, it is a particular meaning of a word (rather than the word itself) that conveys emotion. To mitigate this issue, we present the Emotion Meanings dataset, a novel dataset of 6000 Polish word meanings. The word meanings are derived from the Polish wordnet (plWordNet), a large semantic network interlinking words by means of lexical and conceptual relations. The word meanings were manually rated for valence and arousal, along with a variety of basic emotion categories (anger, disgust, fear, sadness, anticipation, happiness, surprise, and trust). The annotations were found to be highly reliable, as demonstrated by the similarity between data collected in two independent samples: unsupervised (n = 21,317) and supervised (n = 561). Although we found the annotations to be relatively stable for female, male, younger, and older participants, we share both summary data and individual data to enable emotion research on different demographically specific subgroups. The word meanings are further accompanied by the relevant metadata, derived from open-source linguistic resources. Direct mapping to Princeton WordNet makes the dataset suitable for research on multiple languages. Altogether, this dataset provides a versatile resource that can be employed for emotion research in psychology, cognitive science, psycholinguistics, computational linguistics, and natural language processing.
This article presents MANULEX, a Web-accessible database that provides grade-level word frequency lists of nonlemmatized and lemmatized words (48,886 and 23,812 entries, respectively) computed from the 1.9 million words taken from 54 French elementary school readers. Word frequencies are provided for four levels: first grade (G1), second grade (G2), third to fifth grades (G3–5), and all grades (G1–5). The frequencies were computed following the methods described by Carroll, Davies, and Richman (1971) and Zeno, Ivenz, Millard, and Duvvuri (1995), with four statistics at each level (F, overall word frequency; D, index of dispersion across the selected readers; U, estimated frequency per million words; and SFI, standard frequency index). The database also provides the number of letters in the word and syntactic category information. MANULEX is intended to be a useful tool for studying language development through the selection of stimuli based on precise frequency norms. Researchers in artificial intelligence can also use it as a source of information on natural language processing to simulate written language acquisition in children. Finally, it may serve an educational purpose by providing basic vocabulary lists. This article presents MANULEX, 1 the first French linguistic tool that provides grade-based frequency lists of the 1.9 million words found in first-grade, secondgrade, and third- to fifth-grade French elementary school readers. The database contains 48,886 nonlemmatized entries and 23,812 lemmatized entries. It was compiled to supply the French counterpart to such works on the
This paper presents a new corpus of 140 high quality colour images belonging to 14 subcategories and covering a range of naming difficulty. One hundred and six Spanish speakers named the items and provided data for several psycholinguistic variables: age of acquisition, familiarity, manipulability, name agreement, typicality and visual complexity. Furthermore, we also present lexical frequency data derived internet search hits. Apart from the large number of variables evaluated, these stimuli present an important advantage with respect to other comparable image corpora in so far as naming performance in healthy individuals is less prone to ceiling effect problems. Reliability and validity indexes showed that our items display similar psycholinguistic characteristics to those of other corpora. In sum, this set of ecologically valid stimuli provides a useful tool for scientists engaged in cognitive and neuroscience-based research.
This paper revisits the age-of-acquisition (AoA) norms of Kuperman et al. (2012). Three studies were conducted. Study 1 reports a crowdsourcing 'megastudy' obtaining 790,024 estimates from participants with the age they could first read and write 11,074 early acquired words from Kuperman et al. (2012). The study aimed to differentiate between oral language receptive AoA and print-based AoA. The results correlate well with the original estimates, offering, as hypothesized, higher AoAs for reading/writing. These are released as supplements to the original norms. Study 2 explored the potential of large language models (LLMs), specifically GPT-4o, to replicate these crowdsourced AoA estimates. The findings indicated a strong correlation between AI-generated estimates and human judgments, showing the utility of AI in estimating AoA and developing norms for psycholinguistic and educational research in lieu of crowdsourcing. Study 3 leveraged AI to extend estimates to all well-known words in Kuperman et al. (2012) and the English Crowdsourcing Project (ECP). Study 3 also investigated a trained model fine-tuned on 2000 ratings from Kuperman et al. (2012). Fine-tuning increased alignment with human ratings, though comparisons with untrained models suggested that fine-tuning is not essential in English for obtaining useful AoA estimates. Both trained and untrained AI-generated norms correlated highly with human ratings and performed well in accounting for word processing times and accuracy in regressions. Uses and limitations of the AI estimates are discussed. All resources are made available in the Open Science Framework and can be used freely for research and education.
Megastudies and crowdsourcing studies are a rich source of information for word recognition research because they provide processing times for thousands of words. However, the high cost makes it impossible to include all words of interest and all relevant participant groups. This study explores the potential of fine-tuned large language models (LLMs) to generate lexical decision times (RTs) similar to those of humans. Building on recent findings that LLMs can accurately estimate word features, we fine-tuned GPT-4o mini with 3000 words from a megastudy. We then gave the model the task of generating RT estimates for the remaining words in the dataset. Our findings showed a high correlation between AI-generated and observed RTs. We discuss three applications: (1) estimating missing RT data, where AI can fill in gaps for words missing in some megastudies, (2) verifying results of virtual experiments, where AI-generated data can provide an additional layer of validation for results of virtual experiments, and (3) optimizing human data collection, as researchers can run simulations before conducting studies with humans. While AI-generated RTs are not a replacement for human data, they have the potential to increase the flexibility and efficiency of megastudy research.
Objects are commonly described based on their relations to other objects (e.g., associations, semantic similarity, etc.) or their physical features (e.g., birds have wings, feathers, etc.). However, objects can also be described in terms of their actionable properties (i.e., affordances), which reflect interactive relations between actors and objects. While several normed datasets have been developed to categorize various aspects of meaning (e.g., semantic features, cue–target associations, etc.), to date, norms for affordances have not been generated. We address this limitation by developing a set of affordance norms for 2825 concrete nouns. Using an open-response format, we computed affordance strength (AFS; i.e., the probability of an item eliciting a particular action response), affordance proportion (AFP; i.e., the proportion of participants who provided a specific action response), and affordance set size (AFSS; i.e., the total number of unique action responses) for each item. Because our stimuli overlapped with Pexman et al.’s, Behavior Research Methods, 51, 453-466, (2019) body–object interaction norms (BOI), we tested whether AFS, AFP, and AFSS were related to BOI, as objects with more perceived action properties may be viewed as being more interactive. Additionally, we tested the relationship between AFS and AFP and two separate measures of relatedness: cosine similarity (Buchanan et al., Behavior Research Methods, 51, 1849-1863, 2019a, Behavior Research Methods, 51, 1878-1888, 2019b) and forward associative strength (Nelson et al., Behavior Research Methods, Instruments, & Computers, 36(3), 402–407, 2004). All analyses, however, revealed weak relationships between affordance measures and existing semantic norms, suggesting that affordance properties reflect a separate construct.
Semantic priming has been studied for nearly 50 years across various experimental manipulations and theoretical frameworks. Although previous studies provide insight into the cognitive underpinnings of semantic representations, they have suffered from small sample sizes and a lack of linguistic and cultural diversity. In this Registered Report, we measured the size and the variability of the semantic priming effect across 19 languages (n = 25,163 participants analysed) by creating the largest available database of semantic priming values using an adaptive sampling procedure. We found evidence for semantic priming in terms of differences in response latencies between related word-pair conditions and unrelated word-pair conditions. Model comparisons showed that the inclusion of a random intercept for language improved model fit, providing support for variability in semantic priming across languages. This study highlights the robustness and variability of semantic priming across languages and provides a rich, linguistically diverse dataset for further analysis. The Stage 1 protocol for this Registered Report was accepted in principle on 15 July 2022. The protocol, as accepted by the journal, can be found at https://osf.io/u5bp6 (registration) or https://osf.io/q4fjy (preprint version 6, 31 May 2022).
LOFLOC -- Lexic obèrt flechit Occitan (Open Inflected Lexicon of Occitan) Loflòc is a morphological lexicon for Occitan, a Romance language spoken in the south of France and in parts of Italy and Spain. Occitan is not recognized as an official language in France and no standard variety is shared across the linguistic area. To the best of our knowledge, Loflòc is the first publicly available lexicon for Occitan. It contains 680 thousand entries for 57 thousand lemmas. Each entry contains an inflected form, its lemma and its part-of-speech tag according to the Universal Dependencies guidelines. Currently, the lexicon only contains the Lengadocian variety and the classical spelling norm. Nevertheless, it has been shown to be useful even for processing texts from other varieties (for more details, see Vergez-Couret et al., 2024; full reference below).
This paper introduces association norms of German noun compounds as a lexical-semantic resource for cognitive and computational linguistics research on compositionality. Based on an existing database of German noun compounds, we collected human associations to the compounds and their constituents within a web experiment. The current study describes the collection process and a part-of-speech analysis of the association resource. In addition, we demonstrate that the associations provide insight into the semantic properties of the compounds, and perform a case study that predicts the degree of compositionality of the experiment compound nouns, as relying on the norms. Applying a comparatively simple measure of association overlap, we reach a Spearman rank correlation coefficient of rs = 0.5228, p <.000001, when comparing our predictions with human judgements.
The complexity of Chinese orthography has hindered the progress of research in Chinese to the same level of sophistication of that in alphabetic languages such as English. Also, there has been no publicly available resource concerning the decomposition of Chinese characters, which is essential in any attempt to model the cognitive processes of Chinese character recognition. Here we report our construction and analysis of a Chinese lexical database containing the most frequent phonetic compounds decomposed into semantic and phonetic radicals according to Chinese etymology. Each radical was further decomposed into basic stroke patterns according to a Chinese transcription system, Cangjie (Chu, 1979 Laboratory of chu Bong-Foo Retrieved August 25, 2004, from http://www.cbflabs.com/). Other information such as pronunciation and character frequency were also incorporated. We examine the distribution of different types of character, the information skew in phonetic compounds, the relations between subcharacter orthographic units and the pronunciation of the entire character, and the processing implications of these phenomena in terms of universal psycholinguistic principles. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
The present study examines the effect of the goodness of view on the minimal exposure time required to recognize depth-rotated objects. In a previous study, Verfaillie and Boutsen (1995) derived scales of goodness of view, using a new corpus of images of depth-rotated objects. In the present experiment, a subset of this corpus (five views of 56 objects) is used to determine the recognition exposure time for each view, by increasing exposure time across successive presentations until the object is recognized. The results indicate that, for two thirds of the objects, good views are recognized more frequently and have lower recognition exposure times than bad views.
The handbook has three major parts. It begins with an introduction to the topic of corpus linguistics, intended to bring the substantial amount of corpusbased work already done in a variety of research areas to the non-specialist reader's attention. It also provides an outline description of the BNC itself. The bulk of the book however is concerned with the use of the SARA search program. This part consists of a series of detailed task descriptions which (it is hoped) will serve to teach the reader how to use SARA eVectively, and at the same time stimulate his or her interest in using the BNC. There are ten tasks, each of which introduces a new group of features of the software and of the corpus, of roughly increasing complexity. At the end of each task there are suggestions for further related work. The last part of the handbook gives a summary overview of the SARA program's commands and capabilities, intended for reference purposes, details of the main coding schemes used in the corpus, and a select bibliography
This paper presents the Hellenic National (HNC), which is the corpus of Modern Greek developed by the Institute for Language and Speech Processing (ILSP). The presentation describes all stages of the creation of the corpus: collection of the material, tagging and tokenizing, construction of the database and the online implementation which aims at rendering the corpus accessible over Internet to the research community.
The authors report a series of original studies and analyze previous work in the area. "{\ldots} the position is taken that the frequency with which verbal units have been experienced is the fundamental variable responsible for the characteristics which have been used to define meaningfulness." The implications of the frequency hypothesis were tested in 16 experiments. These experiments deal with the effects of the frequency of letters and letter-combinations on serial and paired-associate learning, the effect of "pronunciability" on learning, the effect of frequency on letter-sequence habits, and the difference in the effect of the meaningfulness of a word depending on whether the word is a stimulus or a response.
The International Affective Picture System (IAPS; Center for the Study of Emotion and Attention [CSEA], 1995) is a set of pictures that is widely used in experimental research on emotion and attention. In this study, the normative ratings of a subset of the IAPS were compared with the ratings from a Flemish sample. Eighty Flemish first-year psychology students from the Ghent University (Belgium) rated valence, dominance and arousal for a stratified sample of 60 IAPS pictures. Reliability coefficients indicate that the self-report ratings are internally consistent. Four findings converge upon the idea that the ratings in the Flemish sample are similar to the normative ratings. First, the affective ratings of the pictures in our sample correlated strongly with the North American ratings: .95, .84 and .87, respectively for valence, arousal and dominance. Second, mean valence and arousal ratings of the 60 pictures did not significantly differ between the Flemish and the North American sample. Third, plotting of the valence and arousal ratings in a two-dimensional figure results in a similar boomerang shaped distribution as the North American affective ratings. And fourth, as predicted, this distribution of the valence and arousal ratings shows the same asymmetry between positive and negative pictures as in North American samples.
This paper describes the conversion of ItalwordNet and of a domain WordNet into RDF and their linking to the (L)LOD cloud and to other existing resources. A brief presentation of the resources is given, and the conversion and resulting datasets are described.
This paper introduces the Corpus of Advanced Learner Finnish (LAS2), one of the existing corpora of learner Finnish. The corpus was started at the University of Turku in 2007, and the initial motivation for its collection was to make it possible to deal with novel linguistic challenges posed by academic immigration and to contribute to corpus linguistics, Finnish linguistics and the study of second language acquisition. This paper describes the typological standpoint of the LAS2, its position with respect to other corpora of learner Finnish, the compilation criteria, the annotation applied and the workflow implemented. The corpus consists of three subcorpora of written academic texts of non-native speakers of Finnish. The subcorpora are 1) texts for examination purposes, 2) texts for publishing and graduating purposes, and 3) texts for studying and learning purposes. The informants either study or work in Finnish within academia in Finland. When available, the data has been collected longitudinally. A reference corpus for each subcorpus written by native speakers has also been compiled. Three query tools designed within the framework of the LAS2 are also introduced. These tools enable queries based on any combinations of the linguistic annotation. They can also be used to analyse the typical inner or cotextual variation of any user-specified linguistic node or to create frequency lists of multiword units defined at any level of the annotation. The queries can be limited to a user-specified subset of the data.
This paper presents two database methods for crosslinguistic data collection and comparison: autotypologizing and exemplar-based sampling. Autotypologizing dispenses with a priori defined comparative grids and instead lets structural types emerge inductively through a type list that is constantly updated in response to languages entered in a database. Examplar-based sampling allows identification of a single representative of cross-linguistically heterogeneous structural domains such as case. These two methods are helpful tools in fieldwork. Autotypologizing generates inventories of known types. These inventories update researchers' expectance range for newly encountered types (like published typological surveys, but more dynamically). Examplar-based sampling is useful for writing typological profiles at very early stages of description.
This technical report describes the implementation and use of ChildFreq, a tool for assessing lexical norms of children from one to seven years old. As the name implies, ChildFreq works by extracting word frequencies from a large corpus of child language. These can then be ordered by age or mean length of utterance, and it is also possible to split the data by the children's gender. A query of words to count the frequency of produces both a line chart and a table with more detailed information. The child language data is taken from the English part of the CHILDES database 1 and comprises more than 5,000 transcriptions ,a total of ≈ 3, 500, 000 word tokens. The children's ages range from six months to seven years, with most children being three years old. ChildFreq is freely available online at http://childfreq.sumsar.net .
Word sketches are one-page automatic, corpus-based summaries of a word's grammatical and collocational behaviour. They were first used in the production of the Macmillan English Dictionary and were presented at Euralex 2002. At that point, they only existed for English. Now, we have developed the Sketch Engine, a corpus tool which takes as input a corpus of any language and a corresponding grammar patterns and which generates word sketches for the words of that language. It also generates a thesaurus and 'sketch differences', which specify similarities and differences between near-synonyms. We briefly present a case study investigating applicability of the Sketch Engine to free word-order languages. The results show that word sketches could facilitate lexicographic work in Czech as they have for English.
We present the lexical-semantic net for German "GermaNet" which integrates conceptual ontological information with lexical semantics, within and across word classes. It is compatible with the Princeton WordNet but integrates principlebased modifications on the constructional and organizational level as well as on the level of lexical and conceptual relations. GermaNet includes a new treatment of regular polysemy, artificial concepts and of particle verbs. It furthermore encodes cross-classification and basic syntactic information, constituting an interesting tool in exploring the interaction of syntax and semantics. The development of such a large scale resource is particularly important as German up to now lacks basic online tools for the semantic exploration of very large corpora.
Opinion mining (OM) is a recent subdiscipline at the crossroads of information retrieval and computational linguistics which is concerned not with the topic a document is about, but with the opinions it expresses. OM has a rich set of applications, ranging from tracking users' opinions about products or about political candidates as expressed in online forums, to customer relationship management. In order to aid the extraction of opinions from text, recent research has tried to automatically determine the “PN-polarity” of subjective terms, i.e. identify whether a term that indicates the presence of an opinion has a positive or a negative connotation. Research on determining the “SO-polarity” of terms, i.e. whether a term indeed indicates the presence of an opinion (a subjective term) or not (an objective, or neutral term) has been instead much scarcer. In this paper we describe SentiWordNet, a lexical resource produced by asking an automated classifier ˆ to associate to each synset s of WordNet (version 2.0) a triplet of scores ˆ(s, p) (for p 2 P ={\{}Positive, Negative, Objective{\}}) describing how strongly the terms contained in s enjoy each of the three properties. The method used to develop SentiWordNet is based on the quantitative analysis of the glosses associated to synsets, and on the use of the resulting vectorial term representations for semi-supervised synset classification. The score triplet is derived by combining the results produced by a committee of eight ternary classifiers, all characterized by similar accuracy levels but extremely different classification behaviour. We present the results of evaluating the accuracy of the automatically assigned triplets on a publicly available benchmark. SentiWordNet is freely available for research purposes, and is endowed with a Web-based graphical user interface.
<jats:p> This study uses three archiving efforts at the <jats:italic toggle="yes">New York Times</jats:italic> as a means to analyse the newspaper as an archival object. I study the traditional ‘morgue’ of physical clippings and photos, the <jats:italic toggle="yes">Times’</jats:italic> joint project with Google Cloud to digitize its photo collection, and the <jats:italic toggle="yes">TimesMachine</jats:italic> interactive digital archive, which made scanned editions of printed issues from 1851 to 2002 publicly available online. Based on interviews with staff and analysis of documents describing past and present newspaper archiving practices, it is clear that the digital archive is not a comprehensive copy of an analogue original. There are a significant number of documents stored in physical archives that have not been translated to digital, and whose loss would be detrimental to historians and media scholars alike. Moreover, even the documents that have been scanned and made available as digital objects do not perfectly mirror their analogue equivalents, meaning that information loss is inherent to the digitization process. As active producers of the past for contemporary purposes, these online news archives serve as cultural gatekeepers, actively shaping journalistic practice and reframing current events in reference to the past. </jats:p>
<jats:p>This article discusses the history of a selection of ephemeral adverbial subordinators (Kortmann 1997, 301), i.e., those that mainly originated in the Early Modern English period (16th -17th centuries) but whose subordinating function, however, either became obsolete rather quickly or was subject to further restrictions beyond this period. This phenomenon was particularly frequent in the CCC relations, these are: causality, conditionality and concessivity. The present article analyses a selected number ephemeral conditional subordinators and compares them with the prototypical conditional subordinator if. The methodology is corpus-based, and I examine the data in the Penn Parsed Corpora of Historical English. The examples discussed reveal that ephemeral conditional subordinators are scarce and serve as a clear illustration of the concept of ephemerality in the realm of adverbial subordinators. </jats:p>
<jats:p>The digital era is an era where all information matters can be accessed through digital media. Many digital platforms that can be used to support even the main platform in learning are no exception for Arabic learning platforms. One of the Arabic learning platforms that can be used for the field of linguistics is the Quranic Arabic Corpus Platform. But it is very unfortunate because there are still very few students, Arabic teachers, and also Arabic lecturers who still apply the Quranic Arabic Corpus platform in their teaching. Even though this platform has very many benefits and conveniences that can be obtained by students and lecturers. Research is a descriptive qualitative research literature study by trying to describe an Arabic language learning platform and its use in environmental education. From the results of the study, it was found that the Arabic learning model using The Quranic Arabic Corpus Platform can improve the proficiency of making sentences in Arabic, increase the quantity of Arabic vocabulary, and increase understanding of the position of each word in Arabic so that in short this platform makes it easier for students to learn Arabic grammar. This paper recommends that Arabic learning activities be maximized by using the Quranic Arabic Corpus platform.</jats:p>