1358 norm sets
95 CVCs were rated for pronounceability (p') on a modification of Underwood and Schulz's 9-point scale. The correlation between testretest mean p' ratings of the CVCs was .99. Correlation between mean p' ratings of CVCs by two different groups of raters was .96. Mean p' ratings of the two groups correlated .88 and .95 with Noble's m' ratings of CVCs. Stability of the means in conjunction with the correlations is evidence that the p' scale may prove to be useful in evaluating CVCs used in verbal learning experiments.
This chapter discusses single-word free-association norms for 328 responses from the Connecticut cultural norms for verbal items in categories. There are all sorts of ways in which the associative relation between two or more words can be measured. The stimulus items used in a study described in the chapter are presented with the rank and frequency with which these responses were given to the category name in the Connecticut norms listed beside each word. Stimulus items were selected from only 21 of the 43 categories, and the number of stimuli from each category ranged from 10 to 18 words. The frequency with which a word occurs as a response across the sample population of n = 100 determines the order of its listing in the norms. Thus, out of 100 Ss, 65 respond bird to the stimulus bluejay and this response, as the highest frequency response, is listed first in the associative responses to this stimulus. The normative data are presented in the sequence.
This chapter discusses the history of use of the Kent–Rosanoff list of word association stimuli. It describes the two sets of French word association norms and the conditions under which they were obtained. The difference between speakers of English and of Western European languages in the diversification of their responses is brought out by the use of the rank-frequency function. The chapter presents a comparison of few of the main characteristics of the sets of Kent–Rosanoff norms among countries and across languages. A simple and frequently used measure of similarity of responses among normative studies is to count the number of items for which the primary responses are identical. This method can be used to compare norms in different languages, to the extent that confidence can be placed in the translation. The chapter also presents a classification of the primary responses of French and American students and workers in terms of grammatical classes of stimuli and responses.
This chapter discusses homographs. As an isolated unit, the identical spelling and pronunciation of the word provides no clue as to which meaning is intended. Although interpretation of the isolated unit can be influenced by the relative frequency with which the word is used to denote one rather than the other, meaning that only the surrounding context makes clear the intended meaning. Such words that have the same form but more than one meaning are called homonyms. If the two meanings of the word are represented by the same spelling, the word is a homograph. If the two meanings of the word are represented by the same pronunciation, the word is a homophone. A word can be both a homograph and a homophone. Some words are homographs but not homophones. Because homonyms are heavily dependent on context for their interpretation, they provide useful material for studying the modification of verbal meaning as a function of experimental variations in context.
This chapter discusses the complete German language norms for responses to 100 words from the Kent–Rosanoff association test. The collection of German language norms for the Kent–Rosanoff word association test was undertaken as one phase of a larger project dealing with the manner in which linguistic habits can modify aspects of behavior such as perception, learning, recall, and generalization. The normative project was carried out in the following manner: a translation from English to German was made of the Kent–Rosanoff word association test, the translated test was administered to a normative group of German students, and the results were analyzed to determine the frequency of each response to each stimulus word. The test forms were prepared on two mimeographed sheets, with numbered stimulus words arranged in columns of 25 words each, two columns to a sheet. The 100 stimulus words of the Kent–Rosanoff word association test occur quite frequently in the English language are considered, as a whole, to be emotionally neutral.
This chapter discusses free-association responses to the primary purposes and other responses selected from the Palermo–Jenkins norms. A useful supplement to free-association norms is the additional determination of associations to responses that are elicited on the original test. These supplemental norms increase the number of association hierarchies available to the investigator and also provide information concerning the independent probabilities of chains of words. Free-association responses allow the independent manipulation of associative directionality of pairs of words, for example, A and B, that is, it becomes possible to choose word pairs on the basis of either the A-B or B-A associative strength. Listing of the associative probabilities of other words, given in response to B, can greatly increase the size of the pool of associative triads, that is, A-B-C chains. The original norms can be used to discover A-B-C word chains or word pairs varying in degree of bidirectionality. However, such a procedure identifies only a limited number of usable word pairs or chains.
This chapter discusses the concepts of substitution, context, and association. The elicited sentences cannot give a particularly biased estimate of word contexts for the stimulus words; however, the substitution task elicits far more antonyms than can be expected on the basis of actual contextual similarities of antonyms. The collocational study involves no editing and gives statistics on text frequencies regardless of the grammatical class of the forms. In connected discourse, nouns, verbs, indefinite pronouns, determiners, prepositions, auxiliaries, nominative pronouns, and copulas are most frequent, with adjectives and adverbs relatively infrequent. For all stimulus words, except verbs and gerunds, there is a significant positive relation between association and substitution, that is, the use of semantically and grammatically paradigmatic words as associates. Where there is high commonality in substitutions, responses tend to be paradigmatic. Where there is a high primary in the precontext or high commonality in the five most common postcontexts, the responses tend to be syntagmatic.
This chapter discusses free-association responses to words from the original Kent–Rosanoff word association test obtained from students residing in England and in Australia. It describes the subject populations represented in the present norms and highlights critical aspects of the procedure. The English sample consisted of 200 men and 200 women who were drawn from seven universities located throughout England. Approximately 60{\%} were enrolled in first-year courses in one of the social sciences, while the remainder were studying arts, science, or education. The median age of the subjects was 18 years, 5 months. The Australian samples were drawn from the universities of Sydney and Tasmania. The test was administered individually, using the conventional Kent–Rosanoff procedure. Response latencies were recorded. For the Australian sample, certain other responses were combined: verbs and participles, nouns and adverbs, and variances of nouns, for example, sit and sitting, peace and peaceful, and worm and earthworm.
The study was devised to determine empirically the pronounciability values of 200 nonsense syllables. The rating method of Underwood and Schulz (1960) was employed and 201 college Ss participated. Obtained data were highly reliable in reference to previous research and a marked relationship with m' values was found. Relationships with association value and speed of learning were also determined. It is anticipated that the obtained values will be useful as methodological aids in the design of verbal learning studies.
A matrix is presented of the errors of perception made by 135 men and women listening to three male and three female speakers reading aloud different randomized lists constructed from the letters of the alphabet and the digits 1—9, heard in white noise. Data from a short-term memory (STM) experiment, using simultaneous visual presentation and immediate ordered recall of two selected vocabularies of nine letters and the digits 1-9, are cited as evidence of phonemic confusion between letters and digits in STM. Conrad (1964) established a high correlation between the systematic errors made by Hsteners identifying letters of the alphabet spoken one at a time in white noise (auditory confusions), and those made by subjects when visually presented with strings of alphabetic material for immediate ordered recall. Briefly, items which sounded similar were more probably confused not only in listening, but also in short-term memory (STM). Once this relationship between auditory and STM errors had been demonstrated, it was possible, using Clarke's constant ratio rule (Clarke, 1957), to select from the listening matrix subsets of letters of diifering probabilities of auditory confusion. Conrad {\&} Hull (1964) have shown that such subsets selected for minimal auditory confusability, whether of three or nine items in the vocabulary, are significantly better recalled, after visual sequential presentation, than subsets of letters of high auditory confusability. In STM tasks using both letters and digits, Wickelgren (1965) with auditory pre-sentation, and Hintzman (1965) using visual stimuli, have shown that inter-class errors are also related to phonemic similarity. Definition of the extent and relative confusability of letters with digits has been hampered by the lack of a complete matrix of auditory confusions amongst these stimuli. The provision of such a matrix was the purpose of this study. METHOD AND PEOCEDTTKB
To obtain more complete tables of letter sequences varying in order of approximation to English than those generally available, sequences of zero through fourth-order approximation were computer-generated using tables of single-letter, digram, trigram, and tetragram frequencies. Two sets of tables are presented. One consists of 100 randomly selected 10-letter sequences of each of zero to fourth-order material. The other consists of 40 8-letter sequences of each type, selected with the restriction that no letter appear more than once in the sequence.
L'{\^{a}}ge d'acquisition et la familiarit{\'{e}} d'un mot sont des facteurs d{\'{e}}cisifs pour l'acc{\`{e}}s au lexique, en production comme en perception. Pour favoriser les recherches sur les m{\'{e}}canismes du traitement lexical en Fran{\c{c}}ais, une base de donn{\'{e}}es lexicales a {\'{e}}t{\'{e}} constitu{\'{e}}e pour un corpus de 1225 mots monosyllabiques et bisyllabiques du Fran{\c{c}}ais. Cet article d{\'{e}}crit la m{\'{e}}thode utilis{\'{e}}e pour le recueil des donn{\'{e}}es, l'information brute obtenue, la proc{\'{e}}dure pour traiter cette information brute, et le contenu de la base de donn{\'{e}}es CHACQFAM obtenu apr{\`{e}}s traitement de l'information brute, ainsi que les proc{\'{e}}dures de validation de ce contenu. CHACQFAM est disponible gratuitement sur le site Internet « http://psycholinguistique.unige.ch/ ». Mots-cl{\'{e}
This paper proposes two new methodologies for the placement of series FACTS devices in deregulated electricity market to reduce congestion. Similar to sensitivity factor based method, the proposed methods form a priority list that reduces the solution space. The proposed methodologies are based on the use of LMP differences and congestion rent, respectively. The methods are computationally efficient, since LMPs are the by-product of a security constrained OPF and congestion rent is a function of LMP difference and power flows. The proposed methodologies are tested and validated for locating TCSC in IEEE 14-, IEEE 30- and IEEE 57-bus test systems. Results obtained with the proposed methods are compared with that of the sensitivity method and with exhaustive OPF solutions. The overall objective of FACTS device placement can be either to minimize the total congestion rent or to maximize the social welfare. Results show that the proposed methods are capable of finding the best location for TCSC installation, that suite both objectives. {\textcopyright} 2006 Elsevier B.V. All rights reserved.
Deriving representations of meaning has been a long-standing problem in cognitive psychology and psycholinguistics. The lack of a model for representing semantic and grammatical knowledge has been a handicap in attempting to model the effects of semantic constraints in human syntactic processing. A computational model of high-dimensional context space, the Hyperspace Analogue to Language (HAL), is presented with a series of simulations modelling a variety of human empirical results. HAL learns its representations from the unsupervised processing of 300 million words of conversational text. HAL's high-dimensional context space can be used to (1) provide a basic categorization of semantic and grammatical concepts, (2) model certain aspects of morphological ambiguity in verbs, and (3) provide an account of semantic context effects in syntactic processing. The authors propose that the distributed and contextually derived representations that HAL acquires provide a basis for the subconceptual knowledge that can be used in accounting for a diverse set of cognitive phenomena. ((c) 1997 APA/PsycINFO, all rights reserved)
Six groups, each consisting of 10 Ss, produced strings of 11, 22, or 33 words, under one of two sets of instructions: “structured” instructions which asked S to produce a grammatically acceptable string; “unstructured instructions” which solicited a random, unconstrained string. Each word in S's production was classified into one of six grammatical classes. For the structured instructions the rank ordering of frequency of the six classes was functors (40{\%}), nouns (25{\%}), verbs (13{\%}), adjectives (12{\%}), pronouns (6{\%}), adverbs (3{\%}). For the unstructured instructions the rank ordering of frequency was nouns (68{\%}), adjectives (12{\%}), verbs (11{\%}), functors (4{\%}), adverbs (1{\%}), pronouns (0.55{\%}). The findings were discussed in relation to three questions.
The method of Glaze was used to scale 320 words and paralogs for meaningfulness. One hundred Ss provided data from which three such measures were derived. Employing the most conventional of these measures (percentage of Ss responding in less than 2.5 sec) to select the items to be learned, a validating study demonstrated the usual relationship between association value and speed of learning. Other investigations have employed the materials successfully for purposes of control when the main interest of the experiment was in some other problem. {\textcopyright} 1967 Academic Press Inc.
329 nouns of A-frequency were rated on a 7-point scale for concreteness, specificity, and pronunciability by three different groups of Ss. A fourth group was used to scale the words for m by a 30-sec production method. Means of each word for s, c, and m are presented. The interrelationships among all parameters and the second-order partial correlations between m, s, c and number of letters (L) are discussed. {\textcopyright} 1966 Academic Press Inc.
Each of 464 noun pairs was rated for synonymy on a 7-point scale by 100 college students. The purpose of this study was twofold. First, it was designed to provide memory and psycholinguistic researchers with extensive synonym norms. Second, it was designed to evaluate the effects of encoding order on perceived synonymy. The hypothesis that limited semantic-memory access can cause synonym pairs to be rated as more synonymous in one word-order than in the other was tested by presenting each noun pair to 50 judges in one order and to another 50 judges in the reverse order. Mean synonym ratings (averaged across word orders) ranged from 6.79 to 2.24, thus demonstrating the need for normative data on synonyms. A significant number of noun pairs showed strong directional effects such that perceived synonymy was significantly changed by word encoding order. The practical and theoretical importance of these directional effects are discussed. {\textcopyright} 1979 Academic Press, Inc.
A review of previous word and letter counts in addition to the applications of these counts were reported. A comprehensive count of initial and terminal letters and bigrams was compiled based on the Ku{\v{c}}era and Francis (Computational analysis of present-day American English. Providence: Brown Univ. Press, 1967) corpus of English words. The count included frequency of occurrence and versatility, or number of different words in which letters or bigrams occurred. It was shown how such counts can be used to describe "Englishness" and make predictions as to the information load of letters in words and pseudo words. {\textcopyright} 1982 Academic Press, Inc.
In the first part of this paper bigram frequency counts are given for the first letter- pairs, last letter-pairs and “other” letter-pairs of words of more than three letters. A short discussion of the use and relevance of such tables is given. In the second part, lists of anagram-pairs of words are given for words of length three or more letters, together with approximate percentages of occurrence of such “anagrammatical” words in the English language.
The normal range of reaction in response to any of our stimulus words is largely confined within narrow limits. The frequency tables compiled from test records given by one thousand normal subjects comprise over ninety per cent of the normal range in the average case. With the aid of the frequency tables and the appendix normal reactions, with a very few exceptions, can be sharply distinguished from pathological ones. The separation of pathological reactions from normal ones simplifies the task of their analysis, and makes possible the application of a classification based on objective criteria. By the application of the association test, according to the method here proposed, no sharp distinction can be drawn between mental health and mental disease; a large collection of material shows a gradual and not an abrupt transition from the normal state to pathological states. In dementia pr{\ae}cox, some paranoic conditions, manic-depressive insanity, general paresis, and epileptic dementia the test reveals some characteristic, though not pathognomonic, associational tendencies.
Selecting appropriate stimuli to induce emotional states is essential in affective research. Only a few standardized affective stimulus databases have been created for auditory, language, and visual materials. Numerous studies have extensively employed these databases using both behavioral and neuroimaging methods. However, some limitations of the existing databases have recently been reported, including limited numbers of stimuli in specific categories or poor picture quality of the visual stimuli. In the present article, we introduce the Nencki Affective Picture System (NAPS), which consists of 1,356 realistic, high-quality photographs that are divided into five categories (people, faces, animals, objects, and landscapes). Affective ratings were collected from 204 mostly European participants. The pictures were rated according to the valence, arousal, and approach-avoidance dimensions using computerized bipolar semantic slider scales. Normative ratings for the categories are presented for each dimension. Validation of the ratings was obtained by comparing them to ratings generated using the Self-Assessment Manikin and the International Affective Picture System. In addition, the physical properties of the photographs are reported, including luminance, contrast, and entropy. The new database, with accompanying ratings and image parameters, allows researchers to select a variety of visual stimulus materials specific to their experimental questions of interest. The NAPS system is freely accessible to the scientific community for noncommercial use by request at http://naps.nencki.gov.pl .
A widely agreed-upon feature of spoken word recognition is that multiple lexical candidates in memory are simultaneously activated in parallel when a listener hears a word, and that those candidates compete for recognition (Luce, Goldinger, Auer, {\&} Vitevitch, Perception 62:615-625, 2000; Luce {\&} Pisoni, Ear and Hearing 19:1-36, 1998; McClelland {\&} Elman, Cognitive Psychology 18:1-86, 1986). Because the presence of those competitors influences word recognition, much research has sought to quantify the processes of lexical competition. Metrics that quantify lexical competition continuously are more effective predictors of auditory and visual (lipread) spoken word recognition than are the categorical metrics traditionally used (Feld {\&} Sommers, Speech Communication 53:220-228, 2011; Strand {\&} Sommers, Journal of the Acoustical Society of America 130:1663-1672, 2011). A limitation of the continuous metrics is that they are somewhat computationally cumbersome and require access to existing speech databases. This article describes the Phi-square Lexical Competition Database (Phi-Lex): an online, searchable database that provides access to multiple metrics of auditory and visual (lipread) lexical competition for English words, available at www.juliastrand.com/phi-lex .
In the present study, normative data in Turkish are presented for the 260 color versions of the original Snodgrass and Vanderwart (1980) picture set for the first time. Norms are reported for name and image agreement, age of acquisition (AoA), visual complexity, and conceptual familiarity, together with written word frequency, and numbers of letters and syllables. We collected data from 277 native Turkish adults in a variety of tasks. The results indicated that, whilst several measures displayed language-specific variation, we also reported what seem to be language-independent-that is, universal-measures that show a systematic relationship across several languages. The implications of the reported measures in the domain of psycholinguistic research in Turkish and for wider cross-linguistic comparisons are discussed.
Automatic speech recognition (ASR) has gained significant improvement for major languages such as English and Chinese, partly due to the emergence of deep neural networks (DNN) and large amount of training data. For minority languages, however, the progress is largely behind the main stream. A particularly obstacle is that there are almost no large-scale speech databases for minority languages, and the only few databases are held by some institutes as private properties, far from open and standard, and very few are free. Besides the speech database, phonetic and linguistic resources are also scarce, including phone set, lexicon, and language model. In this paper, we publish a speech database in Kazakh, a major minority language in the western China. Accompanying this database, a full set of phonetic and linguistic resources are also published, by which a full-fledged Kazakh ASR system can be constructed. We will describe the recipe for constructing a baseline system, and report our present results. The resources are free for research institutes and can be obtained by request. The publication is supported by the M2ASR project supported by NSFC, which aims to build multilingual ASR systems for minority languages in China.
For most of the Uralic languages, there is a lack of systematically collected, consequently transcribed and morphologically annotated text corpora. This paper sums up the steps, the preliminary results and the future directions of building a linguistic corpus of some Uralic languages, namely Tundra Nenets, Udmurt, Synya Khanty, and Surgut Khanty. The experiences of building a corpus containing both old and modern, and written and oral data samples are discussed. Principles concerning data collection strategies of languages with different level of vitality and endangerment are discussed. Methodologies and challenges of data processing, and the levels of linguistic annotation are also described in detail.
Orthography-semantics consistency (OSC) is a measure that quantifies the degree of semantic relatedness between a word and its orthographic relatives. OSC is computed as the frequency-weighted average semantic similarity between the meaning of a given word and the meanings of all the words containing that very same orthographic string, as captured by distributional semantic models. We present a resource including optimized estimates of OSC for 15,017 English words. In a series of analyses, we provide a progressive optimization of the OSC variable. We show that computing OSC from word-embeddings models (in place of traditional count models), limiting preprocessing of the corpus used for inducing semantic vectors (in particular, avoiding part-of-speech tagging and lemmatization), and relying on a wider pool of orthographic relatives provide better performance for the measure in a lexical-processing task. We further show that OSC is an important and significant predictor of reaction times in visual word recognition and word naming, one that correlates only weakly with other psycholinguistic variables (e.g., family size, word frequency), indicating that it captures a novel source of variance in lexical access. Finally, some theoretical and methodological implications are discussed of adopting OSC as one of the predictors of reaction times in studies of visual word recognition.
In experimental contexts, affect-related word lists have been widely applied when examining how cognitive processes interact with emotional processes. These lists, however, present limitations when studying the relation between emotion and cognitive processes such as time and number processing because affective words do not inherently contain time or quantity information. Live events, in contrast, are experienced by an observer and therefore inherently carry affect information. Unfortunately, existing life-event lists and inventories have been largely applied within clinical contexts as diagnostic tools, and therefore are not suitable for many experimental contexts because they do not contain a balanced number of reliably positive, negative, and neutral life events. In Experiment 1, we create a standardized affect-related life-events list with 171 positive, negative, and neutral affect-related life events. In Experiment 2, we show that strength of affect and significance of the event are integral dimensions, suggesting that these two features are difficult to separate perceptually. The implications of these findings and some potential future applications of the created life-events list are discussed.
The use of immersive virtual reality as a research tool is rapidly increasing in numerous scientific disciplines. By combining ecological validity with strict experimental control, immersive virtual reality provides the potential to develop and test scientific theories in rich environments that closely resemble everyday settings. This article introduces the first standardized database of colored three-dimensional (3-D) objects that can be used in virtual reality and augmented reality research and applications. The 147 objects have been normed for name agreement, image agreement, familiarity, visual complexity, and corresponding lexical characteristics of the modal object names. The availability of standardized 3-D objects for virtual reality research is important, because reaching valid theoretical conclusions hinges critically on the use of well-controlled experimental stimuli. Sharing standardized 3-D objects across different virtual reality labs will allow for science to move forward more quickly.
Phonotactic probability refers to the frequency with which phonological segments and sequences of phonological segments occur in words in a given language. We describe one method of estimating phonotactic probabilities based on words in American English. These estimates of phonotactic probability have been used in a number of previous studies and are now being made available to other researchers via a Web-based interface. Instructions for using the interface, as well as details regarding how the measures were derived, are provided in the present article. The Phonotactic Probability Calculator can be accessed at http://www.people.ku.edu/-mvitevit/PhonoProbHome.html.
Using appropriate stimuli to evoke emotions is especially important for researching emotion. Psychologists have provided several standardized affective stimulus databases-such as the International Affective Picture System (IAPS) and the Nencki Affective Picture System (NAPS) as visual stimulus databases, as well as the International Affective Digitized Sounds (IADS) and the Montreal Affective Voices as auditory stimulus databases for emotional experiments. However, considering the limitations of the existing auditory stimulus database studies, research using auditory stimuli is relatively limited compared with the studies using visual stimuli. First, the number of sample sounds is limited, making it difficult to equate across emotional conditions and semantic categories. Second, some artificially created materials (music or human voice) may fail to accurately drive the intended emotional processes. Our principal aim was to expand existing auditory affective sample database to sufficiently cover natural sounds. We asked 207 participants to rate 935 sounds (including the sounds from the IADS-2) using the Self-Assessment Manikin (SAM) and three basic-emotion rating scales. The results showed that emotions in sounds can be distinguished on the affective rating scales, and the stability of the evaluations of sounds revealed that we have successfully provided a larger corpus of natural, emotionally evocative auditory stimuli, covering a wide range of semantic categories. Our expanded, standardized sound sample database may promote a wide range of research in auditory systems and the possible interactions with other sensory modalities, encouraging direct reliable comparisons of outcomes from different researchers in the field of psychology.
To simplify the problem of studying how people learn natural language, researchers use the artificial grammar learning (AGL) task. In this task, participants study letter strings constructed according to the rules of an artificial grammar and subsequently attempt to discriminate grammatical from ungrammatical test strings. Although the data from these experiments are usually analyzed by comparing the mean discrimination performance between experimental conditions, this practice discards information about the individual items and participants that could otherwise help uncover the particular features of strings associated with grammaticality judgments. However, feature analysis is tedious to compute, often complicated, and ill-defined in the literature. Moreover, the data violate the assumption of independence underlying standard linear regression models, leading to Type I error inflation. To solve these problems, we present AGSuite, a free Shiny application for researchers studying AGL. The suite's intuitive Web-based user interface allows researchers to generate strings from a database of published grammars, compute feature measures (e.g., Levenshtein distance) for each letter string, and conduct a feature analysis on the strings using linear mixed effects (LME) analyses. The LME analysis solves the inflation of Type I errors that afflicts more common methods of repeated measures regression analysis. Finally, the software can generate a number of graphical representations of the data to support an accurate interpretation of results. We hope the ease and availability of these tools will encourage researchers to take full advantage of item-level variance in their datasets in the study of AGL. We moreover discuss the broader applicability of the tools for researchers looking to conduct feature analysis in any field.
textcopyright} 2018 Psychonomic Society, Inc. In language production research, the latency with which speakers produce a spoken response to a stimulus and the onset and offset times of words in longer utterances are key dependent variables. Measuring these variables automatically often yields partially incorrect results. However, exact measurements through the visual inspection of the recordings are extremely time-consuming. We present AlignTool, an open-source alignment tool that establishes preliminarily the onset and offset times of words and phonemes in spoken utterances using Praat, and subsequently performs a forced alignment of the spoken utterances and their orthographic transcriptions in the automatic speech recognition system MAUS. AlignTool creates a Praat TextGrid file for inspection and manual correction by the user, if necessary. We evaluated AlignTool's performance with recordings of single-word and four-word utterances as well as semi-spontaneous speech. AlignTool performs well with audio signals with an excellent signal-to-noise ratio, requiring virtually no corrections. For audio signals of lesser quality, AlignTool still is highly functional but its results may require more frequent manual corrections. We also found that audio recordings including long silent intervals tended to pose greater difficulties for AlignTool than recordings filled with speech, which AlignTool analyzed well overall. We expect that by semi-automatizing the temporal analysis of complex utterances, AlignTool will open new avenues in language production research.
This article reports the construction of a multimodal annotated database of spoken discourse and co-verbal gestures by native healthy speakers of Cantonese and individuals with language impairment: the Cantonese AphasiaBank. This corpus was established as a foundation for aphasiologists and clinicians to use in designing and conducting research investigations into theoretical and clinical issues related to acquired language disorders in Chinese. Details in terms of the purpose, structure, and levels of annotation of the database (containing part-of-speech-annotated orthographic transcripts with Romanization and the corresponding videos) are described. The discussion presents the challenges of building a spoken database of a language that is not linguistically well-researched and that does not have a standardized written form for many of its lexical items, as well as presenting how these issues were addressed. Most importantly, the article highlights the potential of Cantonese AphasiaBank as a powerful research tool for linguists and psycholinguists.
The Remote Associates Test (RAT) has been used to measure creativity, however few repositories or standardizations of test items exist, like the normative data on 144 items provided by Bowden and Jung-Beeman. comRAT is a computational solver which has been used to solve the compound RAT in linguistic and visual forms, showing correlation to human performance over the normative data provided by Bowden and Jung-Beeman. This paper describes using a variant of comRAT, comRAT-G, to generate and construct a repository of compound RAT items for use in the cognitive psychology and cognitive modelling community. Around 17 million compound Remote Associates Test items are created from nouns alone, aiming to provide control over (i) frequency of occurrence of query items, (ii) answer items, (iii) the probability of coming up with an answer, (iv) keeping one or more query items constant and (v) keeping the answer constant. Queries produced by comRAT-G are evaluated in a study in comparison with queries from the normative dataset of Bowden and Jung-Beeman, showing that comRAT-G queries are similar to the established query set.
Change blindness has been a topic of interest in cognitive sciences for decades. Change detection experiments are frequently used for studying various research topics such as attention and perception. However, creating change detection stimuli is tedious and there is no open repository of such stimuli using natural scenes. We introduce the Change Blindness (CB) Database with object changes in 130 colored images of natural indoor scenes. The size and eccentricity are provided for all the changes as well as reaction time data from a baseline experiment. In addition, we have two specialized satellite databases that are subsets of the 130 images. In one set, changes are seen in rooms or in mirrors in those rooms (Mirror Change Database). In the other, changes occur in a room or out a window (Window Change Database). Both the sets have controlled background, change size, and eccentricity. The CB Database is intended to provide researchers with a stimulus set of natural scenes with defined stimulus parameters that can be used for a wide range of experiments. The CB Database can be found at http://search.bwh.harvard.edu/new/CBDatabase.html .
Words that correspond to a potential sensory experience-concrete words-have long been found to possess a processing advantage over abstract words in various lexical tasks. We collected norms of concreteness for a set of 1,659 French words, together with other psycholinguistic norms that were not available for these words-context availability, emotional valence, and arousal-but which are important if we are to achieve a better understanding of the meaning of concreteness effects. We then investigated the relationships of concreteness with these newly collected variables, together with other psycholinguistic variables that were already available for this set of words (e.g., imageability, age of acquisition, and sensory experience ratings). Finally, thanks to the variety of psychological norms available for this set of words, we decided to test further the embodied account of concreteness effects in visual-word recognition, championed by Kousta, Vigliocco, Vinson, Andrews, and Del Campo (Journal of Experimental Psychology: General, 140, 14-34, 2011). Similarly, we investigated the influences of concreteness in three word recognition tasks-lexical decision, progressive demasking, and word naming-using a multiple regression approach, based on the reaction times available in Chronolex (Ferrand, Brysbaert, Keuleers, New, Bonin, M{\'{e}}ot, Pallier, Frontiers in Psychology, 2; 306, 2011). The norms can be downloaded as supplementary material provided with this article.
This article has two primary aims. The first is to introduce a new Vietnamese text-based corpus. The Corpora of Vietnamese Texts (CVT; Tang, 2006a) consists of approximately 1 million words drawn from newspapers and children's literature, and is available online at www.vnspeechtherapy.com/vi/CVT. The second aim is to investigate potential differences in lexical frequency and distributional characteristics in the CVT on the basis of place of publication (Vietnam or Western countries) and intended audience: adult-directed texts (newspapers) or child-directed texts (children's literature). We found clear differences between adult- and child-directed texts, particularly in the distributional frequencies of pronouns or kinship terms, which were more frequent in children's literature. Within child- and adult-directed texts, lexical characteristics did not differ on the basis of place of publication. Implications of these findings for future research are discussed.
The Berlin Affective Word List (BAWL, V{\~{o}}, Jacobs, {\&} Conrad, Behavior Research Methods, 35, 606-609, 2006) and the BAWL-R (V{\~{o}} et al. in Behavior Research Methods 38, 606-609, 2009) are two commonly used lists to investigate affective properties of German words. The two-dimensional valence and arousal model of affect underlying the BAWL is traditionally contrasted with models describing affect in discrete emotional categories, which, however, are not currently incorporated in the BAWL. In order to allow future studies to investigate affective processing from both perspectives--or to directly compare them--in the present study, we collected data by assigning nouns taken from the BAWL-R to discrete emotion intensities, which in turn allowed the assignment to discrete emotion categories. In the study, we present Discrete Emotion Norms for Nouns-Berlin Affective Word List (DENN-BAWL). Using these ratings and the psycholinguistic indexes from the BAWL-R, the DENN-BAWL allows researchers to design experiments using highly controlled and reliable word material. Data have been archived at www.fu-berlin.de/allgpsy/DENN-BAWL.
We present here emoFinder ( http://usc.es/pcc/emofinder ), a Web-based search engine for Spanish word properties taken from different normative databases. The tool incorporates several subjective word properties for 16,375 distinct words. Although it focuses particularly on normative ratings for emotional dimensions (e.g., valence and arousal) and discrete emotional categories (fear, disgust, anger, happiness, and sadness), it also makes available ratings for other word properties that are known to affect word processing (e.g., concreteness, familiarity, contextual availability, and age of acquisition). The tool provides two main functionalities: Users can search for words matched on specific criteria with regard to the selected properties, or users can obtain the properties for a set of words. The output from emoFinder is highly customizable and can be accessed online or exported to a computer. The tool architecture is easily scalable, so that it can be updated to include word properties from new Spanish normative databases as they become available.
This article provides norms for general taboo, personal taboo, insult, valence, and arousal for 672 Dutch words, including 202 taboo words. Norms were collected using a 7-point Likert scale and based on ratings by psychology students from the Erasmus University Rotterdam in The Netherlands. The sample consisted of 87 psychology students (58 females, 29 males). We obtained high reliability based on split-half analyses. Our norms show high correlations with arousal and valence ratings collected by another Dutch word-norms study (Moors et al.,, Behavior Research Methods, 45, 169-177, 2013). Our results show that the previously found qua-dratic relation (i.e., U-shaped pattern) between valence and arousal also holds when only taboo words are considered. Additionally, words rated high on taboo tended to be rated low on valence, but some words related to sex rated high on both taboo and valence. Words that rated high on taboo rated high on insult, again with the exception of words related to sex many of which rated low on insult. Finally, words rated high on taboo and insult rated high on arousal. The Dutch Taboo Norms (DTN) database is a useful tool for researchers interested in the effects of taboo words on cognitive processing. The data associated with this paper can be accessed via the Open Science Framework (https://osf.io/vk782/).
Faces are widely used as stimuli in various research fields. Interest in emotion-related differences and age-associated changes in the processing of faces is growing. With the aim of systematically varying both expression and age of the face, we created FACES, a database comprising N = 171 naturalistic faces of young, middle-aged, and older women and men. Each face is represented with two sets of six facial expressions (neutrality, sadness, disgust, fear, anger, and happiness), resulting in 2,052 individual images. A total of N = 154 young, middle-aged, and older women and men rated the faces in terms of facial expression and perceived age. With its large age range of faces displaying different expressions, FACES is well suited for investigating developmental and other research questions on emotion, motivation, and cognition, as well as their interactions. Information on using FACES for research purposes can be found at http://faces.mpib-berlin.mpg.de.
Participants judged which of seven facial expressions (neutrality, happiness, anger, sadness, surprise, fear, and disgust) were displayed by a set of 280 faces corresponding to 20 female and 20 male models of the Karolinska Directed Emotional Faces database (Lundqvist, Flykt, {\&} {\"{O}}hman, 1998). Each face was presented under free-viewing conditions (to 63 participants) and also for 25, SO, 100, 250, and 500 msec (to 160 participants), to examine identification thresholds. Measures of identification accuracy, types of errors, and reaction times were obtained for each expression. In general, happy faces were identified more accurately, earlier, and faster than other faces, whereas judgments of fearful faces were the least accurate, the latest, and the slowest. Norms for each face and expression regarding level of identification accuracy, errors, and reaction times may be downloaded from www.psychonomic.org/archive/.
Human locomotion is a fundamental class of events, and manners of locomotion (e.g., how the limbs are used to achieve a change of location) are commonly encoded in language and gesture. To our knowledge, there is no openly accessible database containing normed human locomotion stimuli. Therefore, we introduce the GestuRe and ACtion Exemplar (GRACE) video database, which contains 676 videos of actors performing novel manners of human locomotion (i.e., moving from one location to another in an unusual manner) and videos of a female actor producing iconic gestures that represent these actions. The usefulness of the database was demonstrated across four norming experiments. First, our database contains clear matches and mismatches between iconic gesture videos and action videos. Second, the male actors and female actors whose action videos matched the gestures in the best possible way, perform the same actions in very similar manners and different actions in highly distinct manners. Third, all the actions in the database are distinct from each other. Fourth, adult native English speakers were unable to describe the 26 different actions concisely, indicating that the actions are unusual. This normed stimuli set is useful for experimental psychologists working in the language, gesture, visual perception, categorization, memory, and other related domains.
We report a new multidimensional measure of visual complexity (GraphCom) that captures variability in the complexity of graphs within and across writing systems. We applied the measure to 131 written languages, allowing comparisons of complexity and providing a basis for empirical testing of GraphCom. The measure includes four dimensions whose value in capturing the different visual properties of graphs had been demonstrated in prior reading research—(1) perimetric complexity, sensitive to the ratio of a written form to its surrounding white space (Pelli, Burns, Farell, {\&} Moore-Page, 2006); (2) number of disconnected components, sensitive to discontinuity (Gibson, 1969); (3) number of connected points, sensitive to continuity (Lanthier, Risko, Stolz, {\&} Besner, 2009); and (4) number of simple features, sensitive to the strokes that compose graphs (Wu, Zhou, {\&} Shu, 1999). In our analysis of the complexity of 21,550 graphs, we (a) determined the complexity variation across writing systems along each dimension, (b) examined the relationships among complexity patterns within and across writing systems, and (c) compared the dimensions in their abilities to differentiate the graphs from different writing systems, in order to predict human perceptual judgments (n = 180) of graphs with varying complexity. The results from the computational and experimental comparisons showed that GraphCom provides a measure of graphic complexity that exceeds previous measures in its empirical validation. The measure can be universally applied across writing systems, providing a research tool for studies of reading and writing.
Standardized pictorial stimuli and predictors of successful picture naming are not readily available for Gulf Arabic. On the basis of data obtained from Qatari Arabic, a variety of Gulf Arabic, the present study provides norms for a set of 319 object pictures and a set of 141 action pictures. Norms were collected from healthy speakers, using a picture-naming paradigm and rating tasks. Norms for naming latencies, name agreement, visual complexity, image agreement, imageability, age of acquisition, and familiarity were established. Furthermore, the database includes other intrinsic factors, such as syllable length and phoneme length. It also includes orthographic frequency values (extracted from Aralex; Boudelaa {\&} Marslen-Wilson, 2010). These factors were then examined for their impact on picture-naming latencies in object- and action-naming tasks. The analysis showed that the primary determinants of naming latencies in both nouns and verbs are (in descending order) image agreement, name agreement, familiarity, age of acquisition, and imageability. These results indicate no evidence that noun- and verb-naming processes in Gulf Arabic are influenced in different ways by these variables. This is the first database for Gulf Arabic, and therefore the norms collected from the present study will be of paramount importance for researchers and clinicians working with speakers of this variety of Arabic. Due to the similarity of the Arabic varieties spoken in the Gulf, these different varieties are grouped together under the label "Gulf Arabic" in the literature. The normative databases and the standardized pictures from this study can be downloaded from http://qufaculty.qu.edu.qa/tariq-khwaileh/download-center/ .
Humor ratings are provided for 4,997 English words collected from 821 participants using an online crowd-sourcing platform. Each participant rated 211 words on a scale from 1 (humorless) to 5 (humorous). To provide for comparisons across norms, words were chosen from a set common to a number of previously collected norms (e.g., arousal, valence, dominance, concreteness, age of acquisition, and reaction time). The complete dataset provides researchers with a list of humor ratings and includes information on gender, age, and educational differences. Results of analyses show that the ratings have reliability on a par with previous ratings and are not well predicted by existing norms.
We introduce the Open Affective Standardized Image Set (OASIS), an open-access online stimulus set containing 900 color images depicting a broad spectrum of themes, including humans, animals, objects, and scenes, along with normative ratings on two affective dimensions-valence (i.e., the degree of positive or negative affective response that the image evokes) and arousal (i.e., the intensity of the affective response that the image evokes). The OASIS images were collected from online sources, and valence and arousal ratings were obtained in an online study (total N = 822). The valence and arousal ratings covered much of the circumplex space and were highly reliable and consistent across gender groups. OASIS has four advantages: (a) the stimulus set contains a large number of images in four categories; (b) the data were collected in 2015, and thus OASIS features more current images and reflects more current ratings of valence and arousal than do existing stimulus sets; (c) the OASIS database affords users the ability to interactively explore images by category and ratings; and, most critically, (d) OASIS allows for free use of the images in online and offline research studies, as they are not subject to the copyright restrictions that apply to the International Affective Picture System. The OASIS images, along with normative valence and arousal ratings, are available for download from www.benedekkurdi.com/{\#}oasis or https://db.tt/yYTZYCga .
The use of emoticons and emoji is increasingly popular across a variety of new platforms of online communication. They have also become popular as stimulus materials in scientific research. However, the assumption that emoji/emoticon users' interpretations always correspond to the developers'/researchers' intended meanings might be misleading. This article presents subjective norms of emoji and emoticons provided by everyday users. The Lisbon Emoji and Emoticon Database (LEED) comprises 238 stimuli: 85 emoticons and 153 emoji (collected from iOS, Android, Facebook, and Emojipedia). The sample included 505 Portuguese participants recruited online. Each participant evaluated a random subset of 20 stimuli for seven dimensions: aesthetic appeal, familiarity, visual complexity, concreteness, valence, arousal, and meaningfulness. Participants were additionally asked to attribute a meaning to each stimulus. The norms obtained include quantitative descriptive results (means, standard deviations, and confidence intervals) and a meaning analysis for each stimulus. We also examined the correlations between the dimensions and tested for differences between emoticons and emoji, as well as between the two major operating systems-Android and iOS. The LEED constitutes a readily available normative database (available at www.osf.io/nua4x ) with potential applications to different research domains.
Using the megastudy approach, we report a new database (MEGALEX) of visual and auditory lexical decision times and accuracy rates for tens of thousands of words. We collected visual lexical decision data for 28,466 French words and the same number of pseudowords, and auditory lexical decision data for 17,876 French words and the same number of pseudowords (synthesized tokens were used for the auditory modality). This constitutes the first large-scale database for auditory lexical decision, and the first database to enable a direct comparison of word recognition in different modalities. Different regression analyses were conducted to illustrate potential ways to exploit this megastudy database. First, we compared the proportions of variance accounted for by five word frequency measures. Second, we conducted item-level regression analyses to examine the relative importance of the lexical variables influencing performance in the different modalities (visual and auditory). Finally, we compared the similarities and differences between the two modalities. All data are freely available on our website ( https://sedufau.shinyapps.io/megalex/ ) and are searchable at www.lexique.org , inside the Open Lexique search engine.