1396 norm sets
The paper presents the MULTEXT-East language resources, a multilingual dataset for language engineering research, focused on the morphosyntactic level of linguistic description. The MULTEXT-East dataset includes the morphosyntactic specifications, morphosyntactic lexica, and a parallel corpus, the novel “1984” by George Orwell, which is sentence aligned and contains hand-validated morphosyntactic descriptions and lemmas. The resources are uniformly encoded in XML, using the Text Encoding Initiative Guidelines, TEI P5, and cover 16 languages, mainly from Central and Eastern Europe: Bulgarian, Croatian, Czech, English, Estonian, Hungarian, Macedonian, Persian, Polish, Resian, Romanian, Russian, Serbian, Slovak, Slovene, and Ukrainian. This dataset, unique in terms of languages covered and the wealth of encoding, is extensively documented, and freely available for research purposes. The paper overviews the MULTEXT-East resources by type and language and gives some conclusions and directions for further work.
In this article, we automatically create two large and richly annotated data sets for studying the English dative alternation. With an intrinsic and an extrinsic evaluation, we address the question of whether such data sets that are obtained and enriched automatically are suitable for linguistic research, even if they contain errors. The extrinsic evaluation consists of building logistic regression models with these data sets. We conclude that the automatic approach for detecting instances of the dative alternation still needs human intervention, but that it is indeed possible to annotate the instances with features that are syntactic, semantic and discourse-related in nature. Only the automatic classification of the concreteness of nouns is problematic.
There are two different levels of interoperability for language resources: operational interoperability and conceptual interoperability. The former refers to the standardization of the formal aspects of language resources so that different resources can work together. The latter refers to the standardization of the notional representation of the semantic content of the analysis. This article addresses both issues but focuses on the latter through a description of the annotation and analysis of the International Corpus of English, which is a corpus for the study of English as a global language. The project is parameterised by component, regional sub-corpora and a set of pre-defined textual categories. The one-million-word British component has been constructed, grammatically tagged, and syntactically parsed. This article is first of all a description of steps taken to ensure conformity within the project. These include corpus design, part-of-speech tagging, and syntactic parsing. The article will then present a study that examines the use of adverbial clauses across speech and writing, illustrating the imminent necessity for interoperable analysis of linguistic data.
The Rovereto Emotion and Cooperation Corpus (RECC) is a new resource collected to investigate the relationship between cooperation and emotions in an interactive setting. Previous attempts at collecting corpora to study emotions have shown that this data are often quite difficult to classify and analyse, and coding schemes to analyse emotions are often found not to be reliable. We collected a corpus of task-oriented (MapTask-style) dialogues in Italian, in which the segments of emotional interest are identified using psycho-physiological indexes (Heart Rate and Galvanic Skin Conductance) which are highly reliable. We then annotated these segments in accordance with novel multimodal annotation schemes for cooperation (in terms of effort) and facial expressions (an indicator of emotional state). High agreement was obtained among coders on all the features. The RECC corpus is to our knowledge the first resource with psycho-physiological data aligned with verbal and nonverbal behaviour data.
This paper describes the preparation, recording, analyzing, and evaluation of a new speech corpus for Modern Standard Arabic (MSA). The speech corpus contains a total of 415 sentences recorded by 40 (20 male and 20 female) Arabic native speakers from 11 different Arab countries representing three major regions (Levant, Gulf, and Africa). Three hundred and sixty seven sentences are considered as phonetically rich and balanced, which are used for training Arabic Automatic Speech Recognition (ASR) systems. The rich characteristic is in the sense that it must contain all phonemes of Arabic language, whereas the balanced characteristic is in the sense that it must preserve the phonetic distribution of Arabic language. The remaining 48 sentences are created for testing purposes, which are mostly foreign to the training sentences and there are hardly any similarities in words. In order to evaluate the speech corpus, Arabic ASR systems were developed using the Carnegie Mellon University (CMU) Sphinx 3 tools at both training and testing/decoding levels. The speech engine uses 3-emitting state Hidden Markov Models (HMM) for tri-phone based acoustic models. Based on experimental analysis of about 8 h of training speech data, the acoustic model is best using continuous observation’s probability model of 16 Gaussian mixture distributions and the state distributions were tied to 500 senones. The language model contains uni-grams, bi-grams, and tri-grams. For same speakers with different sentences, Arabic ASR systems obtained average Word Error Rate (WER) of 9.70%. For different speakers with same sentences, Arabic ASR systems obtained average WER of 4.58%, whereas for different speakers with different sentences, Arabic ASR systems obtained average WER of 12.39%.
Emotions are inherent to any human activity, including human–computer interactions, and that is the reason why recognizing emotions expressed in natural language is becoming a key feature for the design of more natural user interfaces. In order to obtain useful corpora for this purpose, the manual classification of texts according to their emotional content has been the technique most commonly used by the research community. The use of corpora is widespread in Natural Language Processing, and the existing corpora annotated with emotions support the development, training and evaluation of systems using this type of data. In this paper we present the development of an annotated corpus oriented to the narrative domain, called EmoTales, which uses two different approaches to represent emotional states: emotional categories and emotional dimensions. The corpus consists of a collection of 1,389 English sentences from 18 different folk tales, annotated by 36 different people. Our model of the corpus development process includes a post-processing stage performed after the annotation of the corpus, in which a reference value for each sentence was chosen by taking into account the tags assigned by annotators and some general knowledge about emotions, which is codified in an ontology. The whole process is presented in detail, and revels significant results regarding the corpus such as inter-annotator agreement, while discussing topics such as how human annotators deal with emotional content when performing their work, and presenting some ideas for the application of this corpus that may inspire the research community to develop new ways to annotate corpora using a large set of emotional tags.
Identifying objects in conversation is a fundamental human capability necessary to achieve efficient collaboration on any real world task. Hence the deepening of our understanding of human referential behaviour is indispensable for the creation of systems that collaborate with humans in a meaningful way. We present the construction of REX-J, a multi-modal Japanese corpus of referring expressions in situated dialogs, based on the collaborative task of solving the Tangram puzzle. This corpus contains 24 dialogs with over 4 h of recordings and over 1,400 referring expressions. We outline the characteristics of the collected data and point out the important differences from previous corpora. The corpus records extra-linguistic information during the interaction (e.g. the position of pieces, the actions on the pieces) in synchronization with the participants’ utterances. This in turn allows us to discuss the importance of creating a unified model of linguistic and extra-linguistic information from a new perspective. Demonstrating the potential uses of this corpus, we present the analysis of a specific type of referring expression (“action-mentioning expression”) as well as the results of research into the generation of demonstrative pronouns. Furthermore, we discuss some perspectives on potential uses of this corpus as well as our planned future work, underlining how it is a valuable addition to the existing databases in the community for the study and modeling of referring expressions in situated dialog.
The Alcohol Language Corpus (ALC) is the first publicly available speech corpus comprising intoxicated and sober speech of 162 female and male German speakers. Recordings are done in the automotive environment to allow for the development of automatic alcohol detection and to ensure a consistent acoustic environment for the alcoholized and the sober recording. The recorded speech covers a variety of contents and speech styles. Breath and blood alcohol concentration measurements are provided for all speakers. A transcription according to SpeechDat/Verbmobil standards and disfluency tagging as well as an automatic phonetic segmentation are part of the corpus. An Emu version of ALC allows easy access to basic speech parameters as well as the us of R for statistical analysis of selected parts of ALC. ALC is available without restriction for scientific or commercial use at the Bavarian Archive for Speech Signals.
Supervised machine learning methods to model word sense often rely on human labelers to provide a single, ground truth label for each word in its context. We examine issues in establishing ground truth word sense labels using a fine-grained sense inventory from WordNet. Our data consist of a sentence corpus of 1,000 sentences: 100 for each of ten moderately polysemous words. Each word was given multiple sense labels—or a multilabel—from trained and untrained annotators. The multilabels give a nuanced representation of the degree of agreement on instances. A suite of assessment metrics is used to analyze the sets of multilabels, such as comparisons of sense distributions across annotators. Our assessment indicates that the general annotation procedure is reliable, but that words differ regarding how reliably annotators can assign WordNet sense labels, independent of the number of senses. We also investigate the performance of an unsupervised machine learning method to infer ground truth labels from various combinations of labels from the trained and untrained annotators. We find tentative support for the hypothesis that performance depends on the quality of the set of multilabels, independent of the number of labelers or their training.
In the context of Systemic Functional Linguistics, Appraisal is a theory describing the types of language utilised in communicating emotion and opinion. Robust automatic analyses of Appraisal could contribute in a number of ways to computational sentiment analysis by: distinguishing various types of evaluation, for example affect, ethics or aesthetics; discriminating between an author’s opinions and the opinions of authors referenced by the author and determining the strength of evaluations. This paper reviews the typology described by Appraisal, presents a methodology for annotating Appraisal, and the use of this to annotate a corpus of book reviews. It discusses an inter-annotator agreement study, and considers instances of systematic disagreement that indicate areas in which Appraisal may be refined or clarified. Although the annotation task is difficult, there are many instances where the annotators agree; these are used to create a gold-standard corpus for future experimentation with Appraisal.
Multilingual text processing is useful because the information content found in different languages is complementary, both regarding facts and opinions. While Information Extraction and other text mining software can, in principle, be developed for many languages, most text analysis tools have only been applied to small sets of languages because the development effort per language is large. Self-training tools obviously alleviate the problem, but even the effort of providing training data and of manually tuning the results is usually considerable. In this paper, we gather insights by various multilingual system developers on how to minimise the effort of developing natural language processing applications for many languages. We also explain the main guidelines underlying our own effort to develop complex text mining software for tens of languages. While these guidelines—most of all: extreme simplicity—can be very restrictive and limiting, we believe to have shown the feasibility of the approach through the development of the Europe Media Monitor (EMM) family of applications (http://emm.newsbrief.eu/overview.html). EMM is a set of complex media monitoring tools that process and analyse up to 100,000 online news articles per day in between twenty and fifty languages. We will also touch upon the kind of language resources that would make it easier for all to develop highly multilingual text mining applications. We will argue that—to achieve this—the most needed resources would be freely available, simple, parallel and uniform multilingual dictionaries, corpora and software tools.
Over recent years, there has been a growing interest in the computational treatment of nominalized Noun Phrases due to the rich semantic information they contain. These Noun Phrases can be understood as verbal paraphrases and, just like them, they can also denote argument and thematic-role relations. This paper presents the methodology followed to annotate the argument structure of deverbal nominalizations in the Spanish AnCora-Es corpus. We focus on the automated annotation process that is mostly based on the semantic information specified in a verbal lexicon but also on the syntactic and semantic information annotated in the corpus. The heuristic rules that make use of this information rely on linguistic assumptions that are also evaluated as we evaluate the reliability of the automated process. The automated annotation was manually checked in order to ensure the accuracy of the final resource. We demonstrate its feasibility (77% F-measure) and show that it facilitates corpus annotation, which is always a time-consuming and costly process. The result is the enrichment of the AnCora-Es corpus with the argument structure and thematic roles of deverbal nominalizations. It is the first Spanish corpus with this kind of information that is freely available.
We present a new online psycholinguistic resource for Greek based on analyses of written corpora combined with text processing technologies developed at the Institute for Language & Speech Processing (ILSP), Greece. The “ILSP PsychoLinguistic Resource” (IPLR) is a freely accessible service via a dedicated web page, at http://speech.ilsp.gr/iplr. IPLR provides analyses of user-submitted letter strings (words and nonwords) as well as frequency tables for important units and conditions such as syllables, bigrams, and neighbors, calculated over two word lists based on printed text corpora and their phonetic transcription. Online tools allow retrieval of words matching user-specified orthographic or phonetic patterns. All results and processing code (in the Python programming language) are freely available for noncommercial educational or research use.
This paper proposes to advance in the current state-of-the-art of automatic Language Resource (LR) building by taking into consideration three elements: (1) the knowledge available in existing LRs, (2) the vast amount of information available from the collaborative paradigm that has emerged from the Web 2.0 and (3) the use of standards to improve interoperability. We present a case study in which a set of LRs for different languages (WordNet for English and Spanish and Parole-Simple-Clips for Italian) are extended with Named Entities (NE) by exploiting Wikipedia and the aforementioned LRs. The practical result is a multilingual NE lexicon connected to these LRs and to two ontologies: SUMO and SIMPLE. Furthermore, the paper addresses an important problem which affects the Computational Linguistics area in the present, interoperability, by making use of the ISO LMF standard to encode this lexicon. The different steps of the procedure (mapping, disambiguation, extraction, NE identification and postprocessing) are comprehensively explained and evaluated. The resulting resource contains 974,567, 137,583 and 125,806 NEs for English, Spanish and Italian respectively. Finally, in order to check the usefulness of the constructed resource, we apply it into a state-of-the-art Question Answering system and evaluate its impact; the NE lexicon improves the system’s accuracy by 28.1%. Compared to previous approaches to build NE repositories, the current proposal represents a step forward in terms of automation, language independence, amount of NEs acquired and richness of the information represented.
Built on the basis of the methods developed for Princeton WordNet and EuroWordNet, Arabic WordNet (AWN) has been an interesting project which combines WordNet structure compliance with Arabic particularities. In this paper, some AWN shortcomings related to coverage and usability are addressed. The use of AWN in question/answering (Q/A) helped us to deeply evaluate the resource from an experience-based perspective. Accordingly, an enrichment of AWN was built by semi-automatically extending its content. Indeed, existing approaches and/or resources developed for other languages were adapted and used for AWN. The experiments conducted in Arabic Q/A have shown an improvement of both AWN coverage as well as usability. Concerning coverage, a great amount of named entities extracted from YAGO were connected with corresponding AWN synsets. Also, a significant number of new verbs and nouns (including Broken Plural forms) were added. In terms of usability, thanks to the use of AWN, the performance for the AWN-based Q/A application registered an overall improvement with respect to the following three measures: accuracy (+9.27 % improvement), mean reciprocal rank (+3.6 improvement) and number of answered questions (+12.79 % improvement).
The English-language Princeton WordNet (PWN) and some wordnets for other languages have been extensively used as lexical–semantic knowledge sources in language technology applications, due to their free availability and their size. The ubiquitousness of PWN-type wordnets tends to overshadow the fact that they represent one out of many possible choices for structuring a lexical–semantic resource, and it could be enlightening to look at a differently structured resource both from the point of view of theoretical–methodological considerations and from the point of view of practical text processing requirements. The resource described here—SALDO—is such a lexical–semantic resource, intended primarily for use in language technology applications, and offering an alternative organization to PWN-style wordnets. We present our work on SALDO, compare it with PWN, and discuss some implications of the differences. We also describe an integrated infrastructure for computational lexical resources where SALDO forms the central component.
The project on the Romanian wordnet has been under continuous development for more than 10 years now. It has been in constant use in many projects and applications which determined, to a large extent, the content and coverage of various lexical domains. The article presents the most recent developments of the Romanian wordnet and offers quantitative data for its current version.
This paper provides a deduction-based approach for automatically classifying compound-internal relations in GermaNet, the German version of the Princeton WordNet for English. More specifically, meronymic relations between simplex and compound nouns provide the necessary input to the deduction patterns that involve different types of compound-internal relations. The scope of these deductions extends to all four meronymic relations modeled in version 6.0 of GermaNet: component, member, substance, and portion. This deduction-based approach provides an effective method for automatically enriching the set of semantic relations included in GermaNet.
The MoveOn speech and noise database was purposely designed and implemented in support of research on spoken dialogue interaction in a motorcycle environment. The distinctiveness of the MoveOn database results from the requirements of the application domain—an information support and operational command and control system for the two-wheel police force—and also from the specifics of the adverse open-air acoustic environment. In this article, we first outline the target application, motivating the database design and purpose, and then report on the implementation details. The main challenges related to the choice of equipment, the organization of recording sessions, and some difficulties that were experienced during this effort, are discussed. We offer a detailed account of the database statistics, the suggested data splits in subsets, and discuss results from automatic speech recognition experiments which illustrate the degree of complexity of the operational environment.
Various areas of research (e.g., memory, metamemory, visual word recognition, associative priming) rely on the careful construction of reliable word lists. ListChecker Pro 1.2 is a computer program that accesses the University of South Florida word association norms (Nelson, McEvoy, {\&} Schreiber, 1998, 2004) to report characteristics of words (e.g., frequency, concreteness), as well as direct and indirect associative relationships (e.g., shared associates, mediators). The present article presents the input requirements, menu options, and output obtained by ListChecker Pro 1.2. In addition, a randomly selected list of words from the associative versus semantic priming literature was submitted to ListChecker Pro 1.2 to demonstrate how seemingly unrelated words can be associated. The zipped file containing the program and database can be downloaded from www.eakinmemorylab.psychology.msstate.edu.
This study provides implicit verb causality norms for a corpus of 305 English verbs. A web-based sentence completion study was conducted, with 96 respondents completing fragments such as "John liked Mary because..." The resulting bias scores are provided as supplementary material in the Psychonomic Society Archive, where we also present lexical and semantic verb features, such as the frequency, semantic class and emotional valence. Our results replicate those of previous studies with much smaller numbers of verbs and respondents. Novel effects of gender and its interaction with verb valence illustrate the type of issues that can be investigated using stable norms for a large number of verbs. The corpus will facilitate future studies in a range of areas, including psycholinguistics and social psychology.
We investigated the linguistic features of temporal cohesion that distinguish variations in temporal coherence. In an analysis of 150 texts, experts rated temporal coherence on three continuous scale measures designed to capture unique representations of time. Coh-Metrix, a computational tool that assesses textual cohesion, correctly predicted the human ratings with five features of temporal cohesion. The correlations between predicted and actual scores were all statistically significant. In a complementary study, we explored the importance of temporal cohesion in characterizing genre. A discriminant function analysis, using Coh-Metrix temporal indices, successfully distinguished the genres of science, history, and narrative texts. The results suggested that history texts are more similar to narrative texts than to science texts in terms of temporal cohesion.
A frequency count of more than 190,000 words of spoken English is presented. The count is based on a published corpus of spontaneous conversation (Svartvik {\&} Quirk, 1980). A brief description of the count is presented, and the correlations between spoken word frequency and a range of other word variables are reported. It is expected that the frequency count will be useful in the interpretation of certain psychological data.
Emotional words are increasingly used in the study of word processing. To elucidate whether the experimental effects obtained with these words are due either to their affective content or to other semantic characteristics, it is necessary to conduct experiments with affectively valenced words obtained from different semantic categories. In the present article, we present affective ratings for 380 Spanish words belonging to three semantic categories: animals, people, and objects. The norms are based on the assessments made by 504 participants, who rated about 47 words either in valence and arousal, by using the Self-Assessment Manikin (Bradley {\&} Lang, Journal of Behavioral Therapy and Experimental Psychiatry, 25, 49-59. 1994), or in concreteness and familiarity. These ratings will help researchers select stimuli for experiments in which both the affective properties of words and their membership to a given semantic category have to be taken into account. The database is available as an online supplement for this article.
Lexical co-occurrence models of semantic memory represent word meaning by vectors in a high-dimensional space. These vectors are derived from word usage, as found in a large corpus of written text. Typically, these models are fully automated, an advantage over models that represent semantics that are based on human judgments (e.g., feature-based models). A common criticism of co-occurrence models is that the representations are not grounded: Concepts exist only relative to each other in the space produced by the model. It has been claimed that feature-based models offer an advantage in this regard. In this article, we take a step toward grounding a cooccurrence model. A feed-forward neural network is trained using back propagation to provide a mapping from co-occurrence vectors to feature norms collected from subjects. We show that this network is able to retrieve the features of a concept from its co-occurrence vector with high accuracy and is able to generalize this ability to produce an appropriate list of features from the co-occurrence vector of a novel concept.
Verb bias, or the tendency of a verb to appear with a certain type of complement, has been employed in psycholinguistic literature as a tool to test competing models of sentence processing. To date, the vast majority of sentence processing research involving verb bias has been conducted almost exclusively with monolingual speakers, and predominantly with monolingual English speakers, despite the fact that most of the world's population is bilingual. To test the generality of competing theories of sentence comprehension, it is important to conduct cross-linguistic studies of sentence processing and to add bilingual data to theories of sentence comprehension. Given this, it is critical for the field to develop verb bias estimates from monolingual speakers of languages other than English and from bilingual populations. We begin to address these issues in two norming studies. Study 1 provides verb bias norming data for 135 Spanish verbs. A second aim of Study 1 was to determine whether verb bias estimates remain stable over time. In Study 2, we asked whether Spanish-English speakers are able to learn verb-specific information, such as verb bias, in their second language. The answer to this question is critical to conducting studies that examine when, during the course of sentence comprehension, bilingual speakers exploit verb information specific to the second language. To facilitate cross-linguistic work, we compared our verb bias results with those provided by monolingual English speakers in a previous norming study conducted by Garnsey, Lotocky, Pearlmutter, and Myers (1997). Our Spanish data demonstrated that individual verbs showed significant similarities in their verb bias across the 3 years of data collection. We also show that bilinguals are able to learn the biases of verbs in their second language, even when immersed in the first language environment. Appendixes A-C, containing the bilingual norms discussed in the article, may be downloaded from http://brm.psychonomic-journals.org/content/supplemental.
The present article introduces SYLLABARIUM, a new Web tool addressing the needs of linguists, psycholinguists, and cognitive scientists who work with Spanish and/or Basque and are interested in retrieving information about several syllable-related parameters. This new online syllabic database allows the user to generate complete lists of Spanish and Basque syllables with information about the syllable frequency. Among other measures, for a given orthographic syllable, SYLLABARIUM provides its number of occurrences (i.e., the type frequency), the summed lexical frequency of the words that contain this syllable (i.e., the token frequency), and the positional distribution of type and token frequencies. The cross-language feature of SYLLABARIUM is of special interest to researchers aiming to explore the influence of the syllable in bilingualism. The Web tool is available at www.bcbl.eu/syllabarium.
Ratings for age of acquisition (AoA) and subjective frequency were collected for the 1,493 monosyllabic French words that were most known to French students. AoA ratings were collected by asking participants to estimate in years the age at which they learned each word. Subjective frequency ratings were collected on a 7-point scale, ranging from never encountered to encountered several times daily. The results were analyzed to address the relationship between AoA and subjective frequency ratings with other psycholinguistic variables (objective frequency, imageability, number of letters, and number of orthographic neighbors). The results showed high reliability ratings with other databases. Supplementary materials for this study may be downloaded from the Psychonomic Society's Archive of Norms, Stimuli, and Data, www.psychonomic.org/archive.
This paper describes a computer search program based on the Medical Research Council Psycholinguistic Database of English words. The program allows words to be extracted from that database according to word length, number of syllables or phonemes, and various psycholinguistic criteria such as frequency of use, imageability, concreteness, meaning, and so forth. Thus it is possible to create, for example, lists of two-syllable words of high and low familiarity. It is also possible to examine properties of given sets of words created by the researcher. Lists of these words with or without their properties and with or without a statistical analysis of those properties may be produced. Particular spellings but not particular phonemes may be searched for.
This article presents affective ratings for 210 British English and Finnish nouns, including taboo words. The norms were collected with 135 native British English and 304 native Finnish speakers, who rated the words according to their emotional valence, emotional charge, offensiveness, concreteness, and familiarity. The ratings between the two languages were found to be strongly correlated. The present ratings were also strongly correlated with the American English emotional valence and arousal ratings available in the Affective Norms for English Words database (Bradley {\&} Lang, 1999) and the Janschewitz (2008) database for taboo words. These ratings will help researchers to select stimulus materials for a wide range of experiments involving both monolingual and bilingual processing of British English and Finnish emotional words. Materials associated with this article may be accessed as an online supplement from http://brm.psychonomic-journals.org/content/supplemental.
Lexical co-occurrence models of semantic memory form representations of the meaning of a word on the basis of the number of times that pairs of words occur near one another in a large body of text. These models offer a distinct advantage over models that require the collection of a large number of judgments from human subjects, since the construction of the representations can be completely automated. Unfortunately, word frequency, a well-known predictor of reaction time in several cognitive tasks, has a strong effect on the co-occurrence counts in a corpus. Two words with high frequency are more likely to occur together purely by chance than are two words that occur very infrequently. In this article, we examine a modification of a successful method for constructing semantic representations from lexical co-occurrence. We show that our new method eliminates the influence of frequency, while still capturing the semantic characteristics of words.
Homographs and homophones have interesting linguistic properties that make them useful in many experiments involving language. To assist researchers in the elicitation of homophones, this paper presents a set of 93 line-drawn pictures of objects with homophonic names and a set of 108 questions with homophonic answers. Statistics are also included for each picture and question: Picture statistics include name-agreement percentages, dominance, and frequency statistics of depicted referents, and picture-naming latencies both with and without study of the picture names. For questions, statistics include answer-agreement percentages, difficulty ratings, dominance, frequency statistics, and naming latencies for 60 of the most consistently answered questions.
Complexity is conventionally defined as the level of detail or intricacy contained within a picture. The study of complexity has received relatively little attention-in part, because of the absence of an acceptable metric. Traditionally, normative ratings of complexity have been based on human judgments. However, this study demonstrates that published norms for visual complexity are biased. Familiarity and learning influence the subjective complexity scores for nonsense shapes, with a significant training x familiarity interaction [F(1,52) = 17.53, p {\textless} .05]. Several image-processing techniques were explored as alternative measures of picture and image complexity. A perimeter detection measure correlates strongly with human judgments of the complexity of line drawings of real-world objects and nonsense shapes and captures some of the processes important in judgments of subjective complexity, while removing the bias due to familiarity effects.
WordGen is an easy-to-use program that uses the CELEX and Lexique lexical databases for word selection and nonword generation in Dutch, English, German, and French. Items can be generated in these four languages, specifying any combination of seven linguistic constraints: number of letters, neighborhood size, frequency, summated position-nonspecific bigram frequency, minimum position-nonspecific bigram f requency, position-specific frequency of the initial and final bigram, and orthographic relatedness. The program also has a module to calculate the respective values of these variables for items that have already been constructed, either with the program or taken from earlier studies. Stimulus queries can be entered through WordGen's graphical user interface or by means of batch files. WordGen is especially useful for (1) Dutch and German item generation, because no such stimulus-selection tool exists for these languages, (2) the generation of nonwords for all four languages, because our program has some important advantages over previous nonword generation approaches, and (3) psycholinguistic experiments on bilingualism, because the possibility of using the same tool for different languages increases the cross-linguistic comparability of the generated item lists. WordGen is free and available at http://expsy.ugent.be/wordgen.htm.
Verb subcategorization frequencies (verb biases) have been widely studied in psycholinguistics and play an important role in human sentence processing. Yet available resources on subcategorization frequencies suffer from limited coverage, limited ecological validity, and divergent coding criteria. Prior estimates of verb transitivity, for example, vary widely with corpus size, coverage, and coding criteria This article provides norming data for 281 verbs of interest to psycholinguistic research, sampled from a corpus of American English, along with a detailed coding manual. We examine the effect on transitivity bias of various coding decisions and methods of computing verb biases.
The present research investigated Internet search engines as a rapid, cost-effective alternative for estimating word frequencies. Frequency estimates for 382 words were obtained and compared across four methods: (1) Internet search engines, (2) the Kucera and Francis (1967) analysis of a traditional linguistic corpus, (3) the CELEX English linguistic database (Baayen, Piepenbrock, {\&} Gulikers, 1995), and (4) participant ratings of familiarity. The results showed that Internet search engines produced frequency estimates that were highly consistent with those reported by Kucera and Francis and those calculated from CELEX, highly consistent across search engines, and very reliable over a 6-month period of time. Additional results suggested that Internet search engines are an excellent option when traditional word frequency analyses do not contain the necessary data (e.g., estimates for forenames and slang). In contrast, participants' familiarity judgments did not correspond well with the more objective estimates of word frequency. Researchers are advised to use search engines with large databases (e.g., AltaVista) to ensure the greatest representativeness of the frequency estimates.
Wordfrequency is one of the strongest determiners of reaction time (RT) in word recognition tasks; it is an important theoretical and methodological variable. The Kucera and Francis (1967) word fre- quency count (derived from the 1-million-word Brown corpus) is used by most investigators concerned with the issue of word frequency. Word frequency estimates from the Brown corpus were compared with those from a 131-million-wordcorpus (the HALcorpus; conversational text gathered from Usenet) in a standard word naming task with 32 subjects. RT was predicted equally well by both corpora for high-frequency words, but the larger corpus provided better predictors for low- and medium-frequency words. Furthermore, the larger corpus provides estimates for 97,261 lexical items; the smaller corpus, for 50,406items.
In the present study, we provide a new technique for the collection of homograph norms that reduces subjectivity in the determination of meaning dominance by allowing participants rather than experimenters to indicate to which meaning or meanings the associates were related. To evaluate the effectiveness of this new technique, a subset of homograph norms were used in three separate experiments, demonstrating that (1) when presented with additional meaning categories, participants classified the associates consistently into the primary and secondary meaning categories; (2) overall, the participants were most familiar with primary meanings, followed by secondary, tertiary, and quaternary meanings; and (3) the meaning categories provided to the participants during norms collection were appropriate, since the two meanings provided for each homograph by the participants were consistent with the original data. Finally, in a fourth experiment, we compared the results of this new technique with a parallel set collected in Australia. The high degree of similarity in the results provides validity for this procedure. The homograph norms discussed in this article may be downloaded from http://brm.psychonomic-journals.org/content/supplemental.
We describe a set of pictorial and auditory stimuli that we have developed for use in word learning tasks in which the participant learns pairings of novel auditory sound patterns (names) with pictorial depictions of novel objects (referents). The pictorial referents are drawings of "space aliens," consisting of images that are variants of 144 different aliens. The auditory names are possible nonwords of English; the stimulus set consists of over 2,500 nonword stimuli recorded in a single voice, with controlled onsets, varying from one to seven syllables in length. The pictorial and nonword stimuli can also serve as independent stimulus sets for purposes other than word learning. The full set of these stimuli may be downloaded from www.psychonomic.org/archive/.
Advances in computational linguistics and discourse processing have made it possible to automate many language- and text-processing mechanisms. We have developed a computer tool called Coh-Metrix, which analyzes texts on over 200 measures of cohesion, language, and readability. Its modules use lexicons, part-of-speech classifiers, syntactic parsers, templates, corpora, latent semantic analysis, and other components that are widely used in computational linguistics. After the user enters an English text, CohMetrix returns measures requested by the user. In addition, a facility allows the user to store the results of these analyses in data files (such as Text, Excel, and SPSS). Standard text readability formulas scale texts on difficulty by relying on word length and sentence length, whereas Coh-Metrix is sensitive to cohesion relations, world knowledge, and language and discourse characteristics.
Associative norms for homographs have been widely used in the study of language processing. A number of sets of these are available, providing the investigator with the opportunity to compare materials collected over a span of years and a range of locations. Words that are homophonic but not homographic have been used to address a variety of questions in memory as well as in language processing. However, a paucity of normative data are available for these materials, especially with respect to responses to the spoken form of the homophone. This article provides such data for a sample of 207 homophones across four different tasks, both visual and auditory, and examines how well the present measures correlate with each other and with those of other investigators. The finding that these measures can account for a considerable proportion of the variance in the lexical decision and naming data from the English Lexicon Project provides an additional demonstration of their utility. The norms from this study are available online in the Psychonomic Society Archive of Norms, Stimuli, and Data, at www .psychonomic.org/archive.
In order to develop an additional measure of availability for the nouns from Paivio, Yuille, and Madigan's (1968) list, we used a CD-ROM version of theOxford English Dictionary (OED) to obtain the number of times a word was used to define other words. This variable was added to Rubin and Friendly's (1986) set of measures for these words. In multiple regression analyses, our measure proved to be a useful predictor of free recall. These results suggest that the OED may be useful for providing additional psycholinguistic measures.
The goal of this study was to test a new technique for assessing vocabulary development. This technique is based on an algorithm for scoring the accuracy of word definitions using a continuous scale (Collins-Thompson {\&} Callan, 2007). In an experiment with adult learners, target words were presented in six different sentence contexts, and the number of informative versus misleading contexts was systematically manipulated. Participants generated a target definition after each sentence, and the definition-scoring algorithm was used to assess the degree of accuracy on each trial. We observed incremental improvements in definition accuracy across trials. Moreover, learning curves were sensitive to the proportion of misleading contexts, the use of spaced versus massed practice, and individual differences, demonstrating the utility of this procedure for capturing specific experimental effects on the trajectory of word learning. We discuss the implications of these results for measurement of meaning, vocabulary assessment, and instructional design.
This study provides normative data on the implicit causality of interpersonal verbs in Spanish. Two experiments were carried out. In Experiment 1, ratings of the implicit causality of 100 verbs classified into four types (agent-patient, agent-evocator, stimulus-experiencer, and experiencer-stimulus) were examined. An offline task was used in which 105 adults and 163 children had to complete sentences containing one verb. Both age and gender effects in the causal biases were examined. In Experiment 2, reading times for sentences containing 60 verbs were analyzed. An online reading task was used in which 34 adults had to read sentences that were both congruent and incongruent with the implicit causality of the verb. The results support the effect of implicit causality in both adults and children, and they support the taxonomy used.
Memory researchers using paired associates have benefited greatly from the Swahili-English norms reported by Nelson and Dunlosky (1994). Given recent increases in the amount and kinds of research using paired associates, however, researchers would now benefit from an expanded set of normative measures for foreign language vocabulary words. We report data for 120 Lithuanian-English word pairs collected from 236 undergraduates. Participants completed three study-test trials and were asked to make metacognitive judgments for each item. We report normative recall performance, recall latencies, and error types for each item across trials, as well as the perceived difficulty of each item on the basis of metacognitive judgments.
A list of role names for future use in research on gender stereotyping was created and evaluated. In two stud- ies, 126 role names were rated with reference to their gender stereotypicality by English-, French-, and German- speaking students of universities in Switzerland (French and German) and in the U.K. (English). Role names were either presented in specific feminine and masculine forms (Study 1) or in the masculine form (generic masculine) only (Study 2). The rankings of the stereotypicality ratings were highly reliable across languages and question- naire versions, but the overall mean of the ratings was less strongly male if participants were also presented with the female versions of the role names and if the latter were presented on the left side of the questionnaires.
A computational analysis of a large British English database was performed and frequencies occurrence of grapheme-phoneme correspondences were obtained. A computer program was implemented, which used these frequencies to predict the probabilities of all possible pronunciations of any given string of graphemes. These results led to a proposal for a quantitative method of measurement of the orthographic depth of different languages.
Acronyms are an idiosyncratic part of our everyday vocabulary. Research in word processing has used acronyms as a tool to answer fundamental questions such as the nature of the word superiority effect (WSE) or which is the best way to account for word-reading processes. In this study, acronym naming was assessed by looking at the influence that a number of variables known to affect mainstream word processing has had in acronym naming. The nature of the effect of these factors on acronym naming was examined using a multilevel regression analysis. First, 146 acronyms were described in terms of their age of acquisition, bigram and trigram frequencies, imageability, number of orthographic neighbors, frequency, orthographic and phonological length, print-to-pronunciation patterns, and voicing characteristics. Naming times were influenced by lexical and sublexical factors, indicating that acronym naming is a complex process affected by more variables than those previously considered.
Although taboo words are used to study emotional memory and attention, no easily accessible normative data are available that compare taboo, emotionally valenced, and emotionally neutral words on the same scales. Frequency, inappropriateness, valence, arousal, and imageability ratings for taboo, emotionally valenced, and emotionally neutral words were made by 78 native-English-speaking college students from a large metropolitan university. The valenced set comprised both positive and negative words, and the emotionally neutral set comprised category-related and category-unrelated words. To account for influences of demand characteristics and personality factors on the ratings, frequency and inappropriateness measures were decomposed into raters' personal reactions to the words versus raters' perceptions of societal reactions to the words (personal use vs. familiarity and offensiveness vs. tabooness, respectively). Although all word sets were rated higher in familiarity and tabooness than in personal use and offensiveness, these differences were most pronounced for the taboo set. In terms of valence, the taboo set was most similar to the negative set, although it yielded higher arousal ratings than did either valenced set. Imageability for the taboo set was comparable to that of both valenced sets. The ratings of each word are presented for all participants as well as for single-sex groups. The inadequacies of the application of normative data to research that uses emotional words and the conceptualization of taboo words as a coherent category are discussed. Materials associated with this article may be accessed at the Psychonomic Society's Archive of Norms, Stimuli, and Data, www.psychonomic.org/archive.
If communality of responses is stable, the relative popularity of responses to the Kent-Rosanoff Word Association Test should remain the same for subjects from young adulthood to advanced age. The Kent-Rosanoff was administered individually to 738 subjects from 18 to 87 years of age from various occupations and from various parts of the country. The results indicate that there is a decrease in the strength of communality accompanied by an increase in variability with the advance in age. ((c) 1997 APA/PsycINFO, all rights reserved)