1396 norm sets
Affective word norms are essential for stimulus control in affective science, experimental psychology, and psycholinguistics, yet Urdu remains underrepresented in established lexical norm resources. The present study developed Urdu valence, arousal, and dominance (VAD) norms for a culturally adapted set of affective words derived from a Pakistan-grounded English VAD lexicon, which was then translated and culturally adapted and normed in Urdu. Using a cross-sectional, laboratory-based design, 240 Pakistani university students and young adults rated one of four stimulus lists. The final dataset comprised 320 Urdu words, with each word receiving 60 valid ratings. For each item, word-level means, standard deviations, valid ns, standard errors, and 95% confidence intervals were computed for valence, arousal, and dominance. The resulting lexicon showed broad coverage across affective space, high internal stability of word-level means across repeated random rater splits, and coherent dimensional associations among valence, arousal, and dominance. Category-based analyses further indicated that the item-construction strategy successfully distributed the final word set across distinct affective regions. The primary contribution of the study is an Urdu VAD norms resource that provides an empirically grounded basis for selecting, matching, and interpreting Urdu verbal stimuli in affective and psycholinguistic research. More broadly, the findings reinforce the importance of establishing affective properties in the target language rather than inferring them from translation alone.
Semantic representations arise from a distillation of multiple sources of information, including sensory, motor, affective, interoceptive, linguistic and cognitive experience. Experience of reward is a highly salient aspect of many human activities, and yet its contribution to semantic processing is not well understood. To address this, the present study took a psycholinguistic approach to measuring and evaluating associations with reward as a facet of word meaning. Behavioural and neurophysiological data suggest that reward processing involves multiple stages and mechanisms. For instance, systems associated with the experience and anticipation of pleasure in response to a reward appear distinct from motivational processes that underlie the pursuit of a stimulus. We sought to collect a novel set of word ratings that capture the full extent of reward-related experience. Initial explorations revealed that reward/pleasure ratings are highly correlated with existing norms of emotional valence. Ratings of association with motivation, however, were only moderately correlated with valence, suggesting they capture distinct semantic information. We therefore conducted a preregistered large-scale study to obtain motivation ratings for 8,601 words. Our analyses suggest these ratings capture aspects of word meaning which are distinct from other semantic dimensions, such as concreteness and valence. Moreover, they explain unique variance in participant performance on lexical, semantic, and recognition memory tasks. We combined motivation and emotional valence ratings to provide a composite measure that might approximate a more general 'reward' construct. However, this did not explain additional variance compared to the individual variables. We discuss the implications of these results for neurocognitive theories of semantics.
The chapter presents and discusses the creation of the spoken corpus LIPS (Lexicon of Spoken Italian by Foreigners). The corpus is based on proficiency exams in Italian as L2 taken at the University for Foreigners of Siena (Università per Stranieri di Siena), and created mainly to enable studies in L2 Italian vocabulary acquisition. The corpus is also organized in a way that facilitates comparisons with the equivalent native speaker corpora. Methodological choices related to transcription norms, lemmatization and part-of-speeh tagging are discussed and motivated. Methods of analysis of native and non-native speakers' lexical richness are discussed as well.
Large Language Models (LLMs) have recently been shown to produce estimates of psycholinguistic norms, such as valence, arousal, or concreteness, for words and multiword expressions, that correlate with human judgments. These estimates are obtained by prompting an LLM, in zero-shot fashion, with a question similar to those used in human studies. Meanwhile, for other norms such as lexical decision time or age of acquisition, LLMs require supervised fine-tuning to obtain results that align with ground-truth values. In this paper, we extend this approach to the previously unstudied features of sentence memorability and reading times, which involve the relationship between multiple words in a sentence-level context. Our results show that via fine-tuning, models can provide estimates that correlate with human-derived norms and exceed the predictive power of interpretable baseline predictors, demonstrating that LLMs contain useful information about sentence-level features. At the same time, our results show very mixed zero-shot and few-shot performance, providing further evidence that care is needed when using LLM-prompting as a proxy for human cognitive measures.
Word frequency is a key variable in psycholinguistics, useful for modeling human familiarity with words even in the era of large language models (LLMs). Frequency in film subtitles has proved to be a particularly good approximation of everyday language exposure. For many languages, however, film subtitles are not easily available, or are overwhelmingly translated from English. We demonstrate that frequencies extracted from carefully processed YouTube subtitles provide an approximation comparable to, and often better than, the best currently available resources. Moreover, they are available for languages for which a high-quality subtitle or speech corpus does not exist. We use YouTube subtitles to construct frequency norms for five diverse languages, Chinese, English, Indonesian, Japanese, and Spanish, and evaluate their correlation with lexical decision time, word familiarity, and lexical complexity. In addition to being strongly correlated with two psycholinguistic variables, a simple linear regression on the new frequencies achieves a new high score on a lexical complexity prediction task in English and Japanese, surpassing both models trained on film subtitle frequencies and the LLM GPT-4. Our code, the frequency lists, fastText word embeddings, and statistical language models are freely available at https://github.com/naist-nlp/tubelex.
"The Romanian-Latin-Hungarian-German Lexicon, printed in Buda, in 1825, is the last and most important of the normative works written by the intellectuals of the Transylvanian School. Considered the first explanatory dictionary of the Romanian language, the lexicon includes, along with orthographic, orthoepic, morphological, lexical norms and etymological information, a special innovation, taken from the scientific lexicography of the time: identifying plants by their scientific name established by Carl von Linnaeus."
Abstract In the current study, Hebrew norms were collected for a set of 320 colored realistic pictures. Interestingly, participants were adult speakers of Hebrew as a first-language (L1) or as a second-language (L2, native Arabic speakers). Thus, both L1 and L2 norming were compiled. For each picture, participants typed its name, and then rated its visual complexity, familiarity, and typicality on scales of 1–7. To establish the predictive utility of the norms, we examined timed picture-naming performance on a subset of 135 items of the normed pictures. Two groups of participants with Hebrew as an L1 (native Hebrew speakers) or as an L2 (native Arabic speakers), were asked to name each picture as quickly and accurately as possible and their reaction times (RT) and accuracy were recorded. Results showed that norms collected from L1 speakers significantly predicted L1 participants’ picture naming RT and accuracy while controlling for objective lexical characteristics (frequency and length), validating the usefulness of the norms. Critically, these same norms were inefficient in predicting L2 picture naming performance. However, norms collected from L2 speakers were significant predictors of L2 picture naming performance. The study, therefore, carries important general implications for L2 production research based on picture naming tasks.
Exploring language usage through frequency analysis in large corpora is a defining feature in most recent work in corpus and computational linguistics. From a psycholinguistic perspective, however, the corpora used in these contributions are often not representative of language usage: they are either domain-specific, limited in size, or extracted from unreliable sources. In an effort to address this limitation, we introduce SubIMDB, a corpus of everyday language spoken text we created which contains over 225 million words. The corpus was extracted from 38,102 subtitles of family, comedy and children movies and series, and is the first sizeable structured corpus of subtitles made available. Our experiments show that word frequency norms extracted from this corpus are more effective than those from well-known norms such as Kucera-Francis, HAL and SUBTLEXus in predicting various psycholinguistic properties of words, such as lexical decision times, familiarity, age of acquisition and simplicity. We also provide evidence that contradict the long-standing assumption that the ideal size for a corpus can be determined solely based on how well its word frequencies correlate with lexical decision times.
Dicos 2020: Occitan Lexicon Online Kathryn Klingebiel Keywords online Occitan lexical resources, Occitan lexicography, paralexicography, Congrès de la Lenga Occitana, Dicod’Òc, collaborative lexicography, crowdsourcing, digitization, multimedia database, Occitan dictionaries, Occitan dialects, nòrma classica, graphie alibertine, decentralization of the norm and of description, Conselh de la Lenga Occitana DICOS 2020 (<http://klingebiel.com/occitan/dicos.html>) continues to broaden its bibliographic documentation of online Occitan lexical resources. In 2016, a short article introduced readers of Tenso 31 to the DICOS site (Klingebiel “Occitan Lexicon Online”). The choice of “lexicon” in the title has proven its suitability, with its broad applicability to any listing of lexemes. DICOS was intended to list virtually everything that could provide access to Occitan lexical resources online. The great surprise from that first round of research was the sheer number of files located, more than 200, in a variety of digital formats. My initial enthusiasm has been tempered by three years’ worth of continued searching and evaluation. Revised and reorganized, DICOS 2020 is presented here with this caveat: while the Occitan lexicon is increasingly easy to explore online, in terms of coverage, relevance, and authenticity it is unevenly served by the internet. DICOS leaves readers free to judge individual resources by the degree to which they conform to the parameters of formal lexicography, that is, analyzing and describing the “semantic, syntagmatic, and paradigmatic relationships within the lexicon of a language” (Wikidiff, s.v. lexicology/lexicography). These parameters are neatly summarized in the materials introducing the Diccionari general de la lenga occitana (DGLO): each entry is intended to specify gender; number; grammatical category; etymology or origin; first attestation; linguistic register; definition (in Occitan); translation (into French, Italian, Castilian, Catalan); a pan-Occitan referent (the most widely-used form across the Occitanophone territory); synonyms and antonyms; related expressions; regional variants; proverbs; literary citations and useful sources. Beyond the traditional parameters of lexicography, DICOS 2020 seamlessly accommodates works of paralexicography and of “lexicographie profane” (“crowdsourced lexicography” as a discipline of “citizen science”), in recognition of the three-way continuum of modern-day lexical resources. [End Page 75] The category of “para-lexicographie grand public” (Margarito 172), often accused of non-professional practices, includes such alternative resources as: glossaries of critical editions and anthologies, vocabularies of famous authors, nineteenth-century compilations of “gasconismes,” children’s picture dictionaries (ima[t]gièrs), listings of proverbs, translation sites, sites for language-learners, popularized listings of toponymy and etymology, and blog listings of “les mots de mon patois,” all found in DICOS 2020. Beyond the scope of DICOS, despite their demonstrable usefulness for lexical documentation, are: synonym, antonym, and rhyme dictionaries; travel phrasebooks; manuals of bon usage; spell-checkers; lexicons for computers; software for manipulating lexical data; and even the French-English-French discussion forum and other translation tools found at <www.reverso.net> and similar sites. Most of these works lying beyond the pale of canonical lexicography, “ouvrages qui occupent une marge floue, mais bien vivante” (Margarito 172), have appeared in response to the needs of electronic media and of crowdsourcing. While they comport certain risks,1 these works offer various advantages, including wide-ranging sources, relatively low cost of revision (as against reprinting), and ease of access and consultation. The modern generation of lexical resources has appeared in three stages: (i) digitization of print works (fr. rétroconversion); (ii) production of online dictionaries and databases; and (iii) creation of collaborative projects. Digitized versions of texts are widely available, e.g., the IEO-Paris’ “Documents per l’estudi de la lenga occitana,” with its more than 120 dictionaries and grammars of Occitan. Digitization has significantly modified the format of many print resources: e.g., the twenty-five volumes of the Französisches etymologisches [End Page 76] Wörterbuch (FEW), for the full Gallo-Romance lexicon, and the Dictionnaire de l’occitan médiéval (DOM), whose seven print fascicles (“a”–“album”) have been reworked into a single searchable online database which now covers “a” through “zyrt.” Selig and Arnold look at changes to the entire DOM infrastructure occurring in the course of digitization. The power of html and of the relational database has been harnessed in the compilation of interlinked...
This repository contains all experimental data, including every respondent's survey, the final data set in Excel or CSV format, and the analysis code in R (norms.R).Paper: https://psyarxiv.com/s2c5h<br>The norms, which are ratings of linguistic stimuli, served a twofold purpose: first, the creation of linguistic stimuli (see also Speed & Majid, 2017), and second, a conceptual replication of Lynott and Connell's (2009, 2013) analyses. In the collection of the ratings, forty-two respondents completed surveys for the properties or the concepts separately. Each word was rated by eight participants on average (see data set), with a minimum of five (e.g., for <em>bevriezend</em>) and a maximum of ten ratings per word (e.g., for <em>donzig</em>). The instructions to participants were similar to those used by Lynott and Connell (2009, 2013), except that we elicited three modalities (auditory, haptic, visual) instead of five.'This is a stimulus validation for a future experiment. The task is to rate how much you experience everyday' [properties/concepts] 'using three different perceptual senses: feeling by touch, hearing and seeing. Please rate every word on each of the three senses, from 0 (not experienced at all with that sense) to 5 (experienced greatly with that sense). If you do not know the meaning of a word, leave it blank.'These norms were validated in an experiment showing that shifts across trials with different dominant modalities incurred semantic processing costs (Bernabeu, Willems, & Louwerse, 2017). All data for that study are available, including a dashboard (in case of downtime of the dashboard site, please see this alternative).The properties and the concepts were analysed separately. Properties were more strongly perceptual than concepts. Distinct relationships also emerged among the modalities, with the visual and haptic modalities being closely related, and the auditory modality being relatively independent (cf. Lynott & Connell's data for English. This ties in with findings that, in conceptual processing, modalities can be collated based on language statistics (Louwerse & Connell, 2011).The norms also served to investigate sound symbolism, which is the relation between the form of words and their meaning. The form of words rests on their sound more than on their visual or tactile properties (at least in spoken language). Therefore, auditory ratings should more reliably predict the lexical properties of words (length, frequency, distinctiveness) than haptic or visual ratings would. Lynott and Connell's (2013) findings were replicated, as auditory ratings were either the best predictor of lexical properties, or yielded an effect that was opposite in polarity to the effects of haptic and visual ratings
While numerous lexical databases provide rating norms for a wide range of words, resources for onomatopoeia remain scarce. Given the pivotal role of onomatopoeia in language development and its potential insights for the relationship between word phonology and word meaning, we introduce the Chinese Onomatopoeia Database (COD), comprising 97 one-character, 380 two-character, 91 three-character, and 183 four-character onomatopoeic words in Chinese (total N = 751). All words were rated by 311 native Chinese speakers for concreteness, imageability, context availability, age of acquisition (AoA), familiarity, semantic transparency, emotional valence, and emotional arousal. We demonstrated high reliability across these measures through Cronbach's alpha, split-half coefficients, and intra-class correlation coefficients (ICCs). Correlation analyses revealed significant associations among these lexical variables, including those between semantic and affective variables. Predictive validity of these variables was also examined using reaction times (RTs) and accuracy (ACC) obtained based on a lexical decision task, which showed that COD variables significantly predicted lexical decision RTs and ACC. Further analyses with two measures, Zipf and logCD, from the Chinese Children's Lexicon of Written Words (CCLOWW; Li et al., 2023) showed that these measures were significantly correlated with all COD variables. Even with the inclusion of CCLOWW-based Zipf or logCD measures in regression models, the COD-based variable still significantly predicted both RTs and ACC. The establishment of the COD not only fills a crucial gap in psycholinguistic resources but also provides a robust tool for future research into the cognitive and developmental underpinnings of language processing.
Sensorimotor information plays a fundamental role incognition. However, datasets of ratings of sensorimotorexperience have generally been restricted to several hundredwords, leading to limited linguistic coverage and reducedstatistical power for more complex analyses. Here, we presentmodality-specific and effector-specific norms for 39,954concepts across six sensory modalities (touch, hearing, smell,taste, vision, and interoception) and five action effectors(mouth/throat, hand/arm, foot/leg, head excluding mouth, andtorso), which were gathered from 4,557 participants whocompleted a total of 32,456 surveys using Amazon'sMechanical Turk platform. The dataset therefore representsone of the largest set of semantic norms currently available.We describe the data collection procedures, provide summarydescriptives of the data set, demonstrate the utility of thenorms in predicting lexical decision times and accuracy, aswell as offering new insights and outlining avenues for futureresearch. Our findings will be of interest to researchers inembodied cognition, cognitive semantics, sensorimotorprocessing, and the psychology of language generally. Thescale of this dataset will also facilitate computationalmodelling and big data approaches to the analysis of languageand conceptual representations.
Free-association norms provide essential empirical data for investigating linguistic, semantic, and cultural phenomena in the cognitive sciences. Although large-scale norms exist for languages such as English, Dutch, Spanish, and Mandarin Chinese, no comparable resource has been available for German. To address this gap, we present free-association norms for 5,877 German cue words as part of the German version of the multilingual Small World of Words (SWOW) project. We describe the data collection procedures, participant characteristics, and our comprehensive preprocessing pipeline before introducing the resulting SWOW-DE data set. Using data from three established psycholinguistic paradigms, we show that SWOW-DE norms robustly predict performance in lexical decision tasks, relatedness judgments, and psycholinguistic word ratings. Furthermore, we demonstrate that SWOW-DE responses compare favorably with existing German resources and provide a preliminary cross-linguistic comparison revealing both shared and language-specific association patterns, highlighting promising directions for future research. Overall, SWOW-DE represents the largest collection of German free associations to date and offers a unique resource for linguistic, psychological, and cross-cultural research.
Research on metaphor has steadily increased over the last decades, as this phenomenon opens a window into a range of processes in language and cognition, from pragmatic inference to abstraction and embodied simulation. At the same time, the demand for rigorously constructed and extensively normed experimental materials increased as well. Here, we present the Figurative Archive, an open database of 997 metaphors in Italian enriched with rating and corpus-based measures (from familiarity to lexical frequency), derived by collecting stimuli used across 11 studies. It includes both everyday and literary metaphors, varying in structure and semantic domains. Dataset validation comprised correlations between familiarity and other measures. The Figurative Archive has several aspects of novelty: it is increased in size compared to previous resources; it includes a novel measure of inclusiveness, to comply with current recommendations for non-discriminatory language use; it is displayed in a web-based interface, with features for a flexible and customized consultation. We provide guidelines for using the Archive in future metaphor studies, in the spirit of open science.
Word Association Norms (WAN) are collections that present stimuli words and the set of their associated responses. The corpus is widely used in diverse areas of expertise. In order to reduce the effort to have a good quality resource that can be reproduced in many languages with minimum sources, a methodology to build Automatic Word Association Norms is proposed (AWAN). The methodology has an input of two simple elements: a) dictionary, and b) pre-processed Word Embeddings. This new kind of WAN is evaluated in two ways: i) learning word embeddings based on the node2vec algorithm and comparing them with human annotated benchmarks, and ii) performing a lexical search for a reverse dictionary. Both evaluations are done in a weighted graph with the AWAN lexical elements. The results showed that the methodology produces good quality AWANs.
Several lexical databases have been developed in both English-speaking countries and other countries, leading to numerous studies using these resources. A prominent example is the English Lexicon Project (ELP; Balota et al., 2007), a large-scale database containing behavioral data on English word processing. The ELP provides data for two main tasks: the lexical decision task (LDT) and the speeded naming task. Among these, the LDT is the most commonly utilized in word-processing research, largely because 1) it is easy to implement, and 2) it can be conducted online with relative ease (Lieber et al., 2014).In the LDT, participants are asked to decide as quickly and accurately as possible whether a visually presented string of letters forms a real word or a non-word. By analyzing the response time from when the string is presented until the participant makes a decision, researchers can evaluate the speed of word access and semantic processing. The LDT has been employed not only to assess word processing efficiency and cognitive load but also to investigate the structure of the mental lexicon and concept representation. For example, researchers have examined the relationship between LDT response times and various word properties, including the frequency effect, where more frequent words are processed faster and reexamined using the LDT data (Brysbaert et al., 2011).Numerous psycholinguistic studies have explored the semantic properties of word recognition using LDT data. Recently, LDT databases have expanded beyond English, with resources available in languages such as Chinese (Tse et al., 2017), French (Ferrand et al., 2017), and Spanish (Aguasvivas et al., 2018), allowing for more efficient research across languages. For example, researchers have tested hypotheses involving grounded cognition and embodied cognition (Barsalou, 2008) in word recognition and explored the relationship between word recognition and sensorimotor information across various languages, e.g., English (Pexman et al., 2019; Sidhu et al., 2014), French (Lalancette et al., 2024), and Spanish (Alonso et al., 2018). They further examined theoretical predictions with large-scale survey data, often using lexical decision task (LDT) reaction times as the dependent variable in regression analyses. While earlier findings have supported these theories by showing consistent trends across languages, recent discussions have highlighted cross-linguistic variability in these effects (Alonso et al., 2018; Lalancette et al., 2024). Such hypothesis testing using a database reduces stimulus bias by incorporating many words (see Dymarska et al., 2023) and enables new discoveries through cross-linguistic comparisons.In Japanese, several databases are available, as will be discussed later. For example, databases exist for attributes such as word imageability (Sakuma et al., 2005) and familiarity (Asahara, 2020), each containing evaluative data for tens of thousands of words. These databases have long been used in various ways, such as serving as control variables in numerous Japanese word recognition studies (e.g., Mizuno and Matsui, 2018; Mochizuki and Ota, 2020, 2024). However, no LDT database currently exists for Japanese, posing a challenge to psycholinguistic research on the Japanese language as a result of limited resources. Of course, lexical decision tasks have been widely used in Japanese word recognition studies (e.g., Kawakami, 2002; Kusunose et al., 2013). However, the number of stimulus words used in these studies is significantly smaller compared to databases such as the ELP (Balota et al., 2007). Furthermore, the data are not always publicly available, which limits their utility as resources. Given the increasing emphasis on cross-linguistic validation—particularly in studies of abstract concepts shaped by language and culture (Dove, 2018)—developing a large-scale Japanese LDT database would not only aid Japanese researchers but also contribute to the broader field. Therefore, this study aimed to construct a Japanese version of LDT database.It is important to note that individual differences in LDT response times exist (e.g., Hawker and Ferraro, 2007; Yates and Slattery, 2019; Lim et al., 2020). To enhance the database, we collected data on participants’ individual characteristics following the LDT. Specifically, participants completed the ENDCOREs, which measures interpersonal communication skills (Fujimoto and Daibo, 2007), and the Japanese version of the Plymouth Sensory Imagery Questionnaire (Psi-Q) (Fukui and Aoki, 2022). The ENDCOREs assesses six dimensions of communication: self-control, expressiveness, comprehension, assertiveness, acceptance of others, and relational adjustment. The Psi-Q evaluates the vividness of mental imagery across sensory modalities (i.e., vision, sound, smell, taste, touch, body, and emotion), capturing individual differences in multisensory imagery.Although we do not hypothesize a direct relationship between these individual difference variables and simple LDT response times (e.g., the higher/lower a score, the slower/faster the response time), they may serve as possible predictors for validating certain content. For instance, the grounded or embodied cognition framework (Barsalou, 2020, 2008) posits that processing words or concepts involves simulating the sensory modalities through which they are acquired. Consistent with this, processing words rich in sensorimotor information tends to be more efficient (Lynott et al., 2020; Siakaluk et al., 2008; Sidhu et al., 2014; Sidhu and Pexman, 2016; Tillotson et al., 2008). Individual differences in sensitivity to sensory and motor modalities may interact with word characteristics and influence LDT performance. Furthermore, the “Words as Social Tools” (WAT) perspective (Borghi and Binkofski, 2014) posits that simulating social and linguistic information is crucial for understanding abstract concepts (Borghi et al., 2019). Therefore, words with a stronger social nature may be processed more efficiently (Diveica et al., 2023), and the interaction between verbal sociality and individual sociality may affect LDT response times. Since ENDCOREs reflect an individual’s communication skills, individuals with high social interaction skills may find it easier to simulate socially relevant words. Consequently, they might be more efficient in processing abstract words with strong social characteristics. While the present study did not specifically examine the relationship between individual differences and LDT response times, future research could benefit from incorporating these variables into the database.This report introduces the Japanese LDT database (JALEX), which incorporates individual differences among participants. The response time and accuracy data can be used for future psycholinguistic studies involving Japanese participants. Additionally, while no hypotheses were tested, future research may explore the role of individual differences as needed.In the development of psycholinguistic norms, approximately 30 to 40 observations per word are typically required (Balota et al., 2007; Ferrand et al., 2017). However, we recruited a relatively large number of participants to account for potential dropouts, as this was an online study, and to develop more reliable norms.Participants were recruited through a crowdsourcing service Yahoo! Crowdsourcing (https://crowdsourcing.yahoo.co.jp/). A total of 2,689 individuals accessed the task. However, 1,037 either did not start, failed to complete the task, or provided no responses. Ultimately, 1,652 participants completed the task. All participants self-reported as native Japanese speakers. Among them, 1,226 were men, 407 were women, two identified as other genders, and 17 chose not to respond. The mean age was 51.07 years (SD = 11.91), with a range from 18 to 85 years. The participants’ highest levels of education were as follows: 26 had completed doctoral programs, 119 had master’s degrees, 1,069 were college graduates, 21 had finished high school, 28 had completed junior high school, and 26 chose not to respond. As detailed below, the words were divided into 38 lists. With 1,652 participants, this resulted in approximately 43 participants per list. To develop JALEX databases, we selected words with semantic properties listed in multiple extant databases (DBs). This approach ensured consistency with previous word recognition studies and supported continuity in future research. We followed a specific selection procedure. First, we used the Word List by Semantic Principles, revised and enlarged edition (WLSP, National Institute for Japanese Language and Linguistics, 2004) as the master list. From this, we selected words that appeared in all eight of the following DBs: the word familiarity DB (Asahara, 2020), an alternate word familiarity DB (Fujita and Kobayashi, 2020), the word frequency DB (Amano and Kondo, 2000), the NINJAL-LWP for TWC word frequency DB (University of Tsukuba et al., 2013), the word difficulty DB (Kajiwara et al., 2020), the imageability DB for visual words (Sakuma et al., 2005), the semantic orientations DB (Takamura et al., 2005), and the abstractness DB for Japanese words (The Social Computing Laboratory, 2021). Following this procedure, we selected 5,736 Japanese words as stimuli. These included 4,977 nouns, 648 verbs, and 111 adjectives.For each word, linguistic characteristics such as orthographic neighborhood size (ONS), phonological neighborhood size (PNS), orthographic Levenshtein distance 20 (OLD20, Yarkoni et al., 2008), the number of letters, and the number of morae (a rhythmic unit of sound) were calculated. The PNS was computed by decomposing the ‘phonetic’ (読み) variable in the WLSP (National Institute for Japanese Language and Linguistics, 2004) by mora and calculating how many words in the WLSP had one mora replaced. Similarly, the ONS was calculated by decomposing the ‘letter (見出し本体)’ variable in the WLSP into individual characters and determining how many words had one letter replaced. OLD20 was calculated using the old20 function in the vwr package (Keuleers, 2013) in R (R Core Team, 2022), based on the ‘letter (見出し本体)’ variable in the WLSP.In addition, non-words were constructed as filler items for the LDT. First, from the WLSP, we excluded words with one mora, words containing spaces, symbols, or particles, homophones, and items with repetitive morae (e.g., ha-ha-ha [ha/ha/ha]), as these could not be transformed into non-words using the procedure described below. The remaining items were then decomposed into morae, and each mora was randomly shuffled. If the resulting item was not found in the WLSP, it was considered a non-word candidate. This process yielded 63,305 non-word candidates, from which we randomly selected 5,736 to serve as fillers for the LDT. The authors reviewed these candidates, and those deemed too similar to real words were replaced with different non-word candidates. All non-word stimuli are available for reference on Open Science Framework (OSF).The words and non-words were randomly divided into 38 lists, each containing 150 or 151 words (150 × 2 + 151 × 36 = 5,736) with an equal number of non-words.The LDT task was conducted online, and participants accessed the LDT program via their own PCs. The program was created using PsychoPy (Peirce et al., 2019) and hosted on Pavlovia (https://pavlovia.org/). After obtaining informed consent from the participants, they were instructed to begin the task. In the LDT, a blank screen appeared for 200 ms, followed by a fixation point in the center of the screen for 300 ms. A string of characters was then presented, and participants had to decide as quickly and accurately as possible whether the string represented a real Japanese word. The string remained on the screen until a response was made or for up to 2,000 ms. Participants pressed the ‘L’ key for words and the ‘S’ key for non-words. If the response was correct, the task proceeded to the next trial; if incorrect, a feedback message ("Wrong") appeared in red. If no response was given within 2,000 ms, the feedback message ("Too late") was displayed in red for 300 ms. Words and non-words were presented in random order. Participants completed 20 practice trials before starting the actual task. The practice trials used different stimuli from those in the actual task.During the task, participants were allowed to take a break for a maximum of 60 seconds between the 100th and 200th trials. During the break, their percentage of correct answers was displayed to encourage them to continue. Upon completing the LDT, participants answered the ENDCOREs (Fujimoto and Daibo, 2007) and Psi-Q (Fukui and Aoki, 2022) questionnaires. Additionally, they provided demographic information, including gender, age, dominant hand, highest level of education, and native language. Data were collected on May 20 and May 21, 2024. We calculated the accuracy rate for each participant, and the lowest percentage of correct responses exceeded 75%. Since no participants demonstrated an exceptionally low accuracy rate, data from all participants were retained for analysis.The procedure for processing the response time data followed that employed in ELP (Balota et al., 2007). First, we extracted only correct trials, where the "L" key was pressed for word stimuli, and excluded any trials with response times below 200 ms. Second, we removed trials that deviated by ±3 SD from the participant’s mean response time. This resulted in the exclusion of 1.96% of trials as outliers.The distribution of response times averaged by item is showed in Figure 1. We presented the partial correlations with existing DB variables referenced in stimulus selection to examine the convergent validity of the response time data (Figure 2). These findings confirmed the phenomena predicted in prior studies. Specifically, we confirmed the frequency effect (Rubenstein et al., 1971) and familiarity effect (Connine et al., 1990), where lexical decision times decrease as word frequency and familiarity increase. We also observed the imageability effect (Balota et al., 2004), where higher imageability leads to faster lexical decisions, and the orthographic similarity effect (Yarkoni et al., 2008), where greater Levenshtein distance results in longer response times. While few studies have reported simple or partial correlations with these variables in Japanese, several experimental studies using Japanese words as stimuli have observed effects similar to those identified in the present study. For instance, Japanese word recognition research has reported faster word processing for words with higher imageability (Ogawa and Nittono, 2018) and higher frequency (Mizuno and Matsui, 2015). The relationship between response time and Kajiwara’s (2020) difficulty rating has yet to be investigated. However, it is reasonable to predict that more difficult words would require longer processing times for comprehension. The current analysis identified a slight positive correlation between word difficulty and response time. These findings suggest that JALEX is valid to a considerable extent.The partial correlations between familiarity, semantic orientation, abstractness, and response time were significant, but the effects were small. Of these, the zero-order correlation for familiarity was r = -.42, suggesting that higher familiarity facilitates responses when not adjusted for covariates. Zero-order correlations for abstractness revealed a small effect (r =.11), indicating that processing was slightly suppressed for more abstract words, consistent with the representativeness effect (Cortese and Balota, 2012). This study also found that words with high ONS had shorter lexical decision times. The results showed that high ONS words took less time to judge than low ONS words when using Kanji words (Mizuno and Matsui, 2014), which is consistent with the current results. However, when using Katakana words, the inhibitory effect was observed, indicating that low ONS words took less time to judge than high ONS words (Kawakami, 2002). Furthermore, an interaction between ONS and PNS has also been observed in lexical decision performance for katakana words (Hino et al., 2011). The difference in these results may be caused by the limited number of words used in the experiment and the factors of the orthographic form. In studies using word norms, it is particularly important to consider the extent of word coverage and the absence of bias (Dymarska et al., 2023). The failure to replicate the effects observed in previous studies in the present analysis of a relatively large database may be attributed to biases in the stimulus sets used in those studies, which could have significantly influenced their results. Future research should assess the reproducibility of findings from previous studies by leveraging large databases, such as JALEX, and conducting comprehensive analyses.In the present study, the imageability effect (Balota et al., 2004) was replicated even after controlling for linguistic statistical variables, such as the frequency of neighboring words. The effects of psycholinguistic variables, such as the imageability effect, are often discussed in relation to semantic richness (Pexman et al., 2013). Semantic richness refers to the idea that words associated with more semantic information have richer semantic representations, enabling them to be processed more quickly and accurately. In other words, our study replicates in Japanese the finding that the ease of forming a mental image is an important semantic variable in word representations. As discussed in the introduction, the relationship between sensorimotor information and word recognition is explained by the concept of semantic richness—specifically, the richness of the semantic dimension of sensorimotor information facilitates word recognition. In future studies, it will be important to investigate the nature of concept representations by examining psycholinguistic variables influencing word recognition beyond imageability.Furthermore, this study is the first DB of LDT to include individual difference variables for respondents, paving the way for future research on individual differences using JALEX. In word recognition research, it has been observed that certain words exhibit significant individual differences and high variability in ratings of psychological variables (Paisios et al., 2023). A key limitation of the previous DB of LDT is that they did not provide individual difference data for participants, making them unsuitable for studying individual differences in words with high variability in ratings among individuals. Future research using JALEX is expected to refine further grounded cognition theory (Barsalou, 2020, 2008) and advance WAT theory (Borghi and Binkofski, 2014), particularly by promoting individual difference studies on the simulation of sensorimotor information and those related to social communication.This database represents the most comprehensive dataset on the efficiency of Japanese visual word processing and stands as a powerful resource for future research in psychology and linguistics. A unique feature of this database is its inclusion of individual difference variables for participants, allowing researchers to analyze these differences in future studies.However, it is important to note that some words in the dataset had lower accuracy rates. For example, at least 15 items had a correct response rate below 70%, with fewer than 20 observations. Items with fewer observations may exhibit lower reliability and reproducibility compared to others. While we did not exclude these items in the current analysis, researchers should be mindful of their presence when using the database.All data reported in this study can be found in the OSF Repository (https://osf.io/qr2sg). Information from existing databases used for validation cannot be included in the data resources of this study as a result of copyright, however, such information is available in the literature.
Extending psycholinguistic research into the lexical representation of two-kanji compound words within the Japanese mental lexicon (Joyce, 2002, 2004), this paper reports on a large-scale word association survey for basic Japanese vocabulary. The database of word association norms, which is being compiled from various survey formats including a web-based version of the survey, supplements existing databases concerning the lexical features of Japanese vocabulary (Amano & Kondo, 1999; Yokoy m a a, Sasahara, Nozaki & Long, 1998), such as familiarity ratings and frequency counts, which are essential for cognitive science research. A particularly promising application of the word association norms data, however, is the creation of lexical association network maps that capture important proprieties of words and their interconnectivity. These maps complement other approaches that attempt to tap into aspects of lexical knowledge, such as WordNet, thesauri, ontologies, and collocation data, while avoiding some of their problems. There are also direct and interesting lexicographical and Japanese language learning applications of the
Résumé Cet article présente des normes d’imageabilité (ou valeurs d’imagerie) pour un ensemble de 1493 mots. Des analyses statistiques réalisées sur ces normes révèlent une fidélité élevée. Les scores d’imageabilité se révèlent par ailleurs assez modestement corrélés avec d’autres variables psycholinguistiques (par ex., fréquences lexicales, âge d’acquisition). Des analyses restreintes à un sous-échantillon de mots en français, ainsi que d’autres sur des mots normés pour l’anglais, révèlent que le nombre de traits sémantiques est modérément positivement corrélé aux scores d’imageabilité, contrairement à l’hypothèse selon laquelle la richesse sémantique est adéquatement indexée par l’imageabilité.
Résumé Les mots homonymes (par exemple, « avocat ») sont largement utilisés dans des expériences en psychologie cognitive afin d'étudier le traitement du langage, la levée des ambiguïtés lexicales et l'organisation en mémoire des représentations lexicales et sémantiques. Ces expériences requièrent le contrôle des relations associatives entre différents stimulus et des fréquences relatives des différentes acceptions des mots homonymes. L'objectif principal des normes d'associations verbales que nous présentons est de permettre aux chercheurs de réaliser de tels contrôles. Chacun des 162 items ambigus a été présenté à 100 sujets dans une tâche d'association libre. La totalité des réponses est présentée, ainsi que la fréquence relative des acceptions estimées à partir de ces normes. Mots clés: normes d'association libre, ambiguïté lexicale, homonymie, fréquence relative des acceptions.
Tangram pictures are abstract pictures which may be used as stimuli in various fields of experimental psychology and are often used in the field of dialogue psychology. The present study provides the first norms for a set of 332 tangram pictures. These pictures were standardized on a set of variables classically used in the literature on cognitive processes, such as visual perception, language, and memory: name agreement, image agreement, familiarity, visual complexity, image variability, and age of acquisition. Furthermore, norms for concreteness were also provided owing to the influence of this variable on the processes involved in lexical production. Correlational analyses on all variables were performed on the data collected from French native speakers. This new set of standardized pictures constitutes a reliable database for researchers when they select tangram pictures. Given the abstract nature of tangram pictures, this paper also discusses the similarities and differences with the literature on line drawings, and highlights their value for dialogue psychology studies, for psycholinguistics studies, and for cognitive psychology in general.
In this paper, two word association (WA) studies are presented in support of recent arguments against the use of native-speaker (NS) norms in WA research. In Study 1, first-language (L1) and second-language (L2) WA norms lists were developed and compared to learner responses as a means of measuring L2 proficiency. The results showed that L2 norms provided a more sensitive measure of L2 lexical development than did traditional NS norms. Study 2 was designed to test the utility of native norms databases in predicting the primary WA responses of Japanese learners to high-frequency English cues. With the exception of only extremely frequent cues, it was shown that native norms were not successful in predicting learner responses. The results of both studies are discussed in terms of cultural and linguistic differences, geographic distance, and dissimilarities in word knowledge between respondent populations. Finally, a proposal is made for the construction of a Japanese WA database of English responses (J-WADE). The methods by which it will be developed, key features, and employment in future research are outlined.
Several norms of psycholinguistic features of Chinese characters exist in Mandarin Chinese, but only a few are available in Cantonese or in the traditional script, and none includes semantic radical transparency ratings. This study presents subjective ratings of age-of-acquisition (AoA), familiarity, imageability, concreteness, and semantic radical transparency in 4376 Chinese characters. The single Chinese characters were rated individually on the five dimensions by 20 native Cantonese speakers in Hong Kong to form the Hong Kong Chinese Character Psycholinguistic Norms (HKCCPN). The split-half reliability and intra-class correlations testified to the high internal reliability of the ratings. Their convergent and discriminant patterns in relations to other psycholinguistic measures echoed previous findings reported on Chinese. There were high correlations for semantic radical transparency, imageability and concreteness, and moderate-to-high correlations for AoA and familiarity among subsets of items that had been collected in previous studies. Concurrent validity analyses showed convergence in predicting behavioral response times in various tasks (lexical decision, naming, and writing-to-dictation) when compared with other Chinese character databases. High predictive validity was shown in writing-to-dictation data from an independent sample of 20 native Cantonese speakers. Several objective psycholinguistic measures (character frequency, stroke number, number of words formed, number of homophones and number of meanings) were included in this database to facilitate its use. These new ratings extend the currently available norms in language and reading research in Cantonese Chinese for researchers, clinicians, and educators, as well as provide them with a wider choice of stimuli.
In the domain of cognitive studies on the lexico-semantic representational system, one of the most important means of ensuring effective experimental designs is using ecological stimulus sets accompanied by normative data on the most relevant variables affecting the processing of their items. In the context of image sets, color photographs are particularly suited to this purpose as they reduce the difficulty of visual decoding processes that may emerge with traditional image sets of line drawings. This is especially so in clinical populations. In this study we provide Italian norms for a set of 357 high quality image-items belonging to 23 semantic subcategories from the Moreno-Martínez and Montoro database. Data from several variables affecting image processing were collected from a sample of 255 Italian-speaking participants: age of acquisition, familiarity, lexical frequency, manipulability, name agreement, typicality and visual complexity. Lexical frequency data were derived from the CoLFIS corpus. Furthermore, we collected data on image oral naming latencies to explore how the variance in these latencies could be explained by these critical variables. Multiple regression analyses on the naming latencies show classical psycholinguistic phenomena, such as the effects of age of acquisition and name agreement. In addition, manipulability was also a significant predictor. The described Italian normative data and naming latencies are available for download as supplementary material.
The aim of this research is to present a Spanish Word Association Norms (WAN) database of concrete nouns. The database includes 234 stimulus words (SWs) and 67,622 response words (RWs) provided by 478 young Mexican adults. Eight different measures were calculated to quantitatively analyze word-word relationships: 1) Associative strength of the first associate, 2) Associative strength of the second associate, 3) Sum of associative strength of first two associates, 4) Difference in associative strength between first two associates, 5) Number of different associates, 6) Blank responses, 7) Idiosyncratic responses, and 8) Cue validity of the first associate. The resulting database is an important contribution given that there are no published word association norms for Mexican Spanish. The results of this study are an important resource for future research regarding lexical networks, priming effects, semantic memory, among others.
This study investigated the lexical-semantic space organized by the semantic and affective features of Indonesian words and their relationship with gender and cultural aspects. We recruited 1,402 participants who were native speakers of Indonesian to rate affective and lexico-semantic properties of 1,490 Indonesian words. Valence, Arousal, Dominance, Predictability, Subjective Frequency, and Concreteness ratings were collected for each word from at least 52 people. We explored cultural differences between American English ANEW (affective norms for English words), Spanish ANEW, and the new Indonesian inventory [called CEFI (concreteness, emotion, and subjective frequency norms for Indonesian words)]. We found functional relationships between the affective dimensions that were similar across languages, but also cultural differences dependent on gender.
This paper introduces a novel collection of word embeddings, numerical representations of lexical semantics, in 55 languages, trained on a large corpus of pseudo-conversational speech transcriptions from television shows and movies. The embeddings were trained on the OpenSubtitles corpus using the fastText implementation of the skipgram algorithm. Performance comparable with (and in some cases exceeding) embeddings trained on non-conversational (Wikipedia) text is reported on standard benchmark evaluation datasets. A novel evaluation method of particular relevance to psycholinguists is also introduced: prediction of experimental lexical norms in multiple languages. The models, as well as code for reproducing the models and all analyses reported in this paper (implemented as a user-friendly Python package), are freely available at: https://github.com/jvparidon/subs2vec.
Project on Linguistic Analysis, Berkeley.
This paper gives a brief survey of the Saarbrücken project on Old Icelandic legal texts, sponsored by the German Research Society, within the Special Research Area “Computer linguistics.” The project's main points of interest are (1) producing adequate machine-readable versions and parsed indices of all legal texts in Old Icelandic, (2) graphemic studies of legal manuscripts, and (3) studies of the distribution and valence of the verbs in those texts. A proposal for encoding Old Norse/Old Icelandic demonstrates how texts of different standards (normalized, diplomatic, graphetic) can be encoded as compatibly as possible. The description of a combined normalization-lemmatization process reveals that even little normalization in a diplomatic text will save much manual parsing.
The Opera del Vocabolario Italiano was given a mandate in 1964 to create a Historical Dictionary of the Italian Language. The main objective was to provide a tool which would give vital information on the development of the Italian language from its origins to the present day. In 1986 the Center incorporated modern computer technology into the project and this led to a series of decisions which affected the nature and the outcome of the project. This article traces the development of the project, and describes both hardware and software systems used, as well as the nature of the relational database being created and its linguistic applications.
The Corpus dei Manoscritti Copti Letterari is a project whose original aim was to reconstruct the Coptic codices from the White Monastery in Upper Egypt. The project was later expanded to include all Coptic literature. In 1980 a new project was launched to transfer the data into machine-readable form and make the information available, in as generic a format as possible, to scholars throughout the world.
CLIPON is an acronym for Concordanze della Lingua Italiana Poetica dell'Otto/Novecento. The aim of the project described here is to produce lexicons and lemmatized concordances of the literary Italian language of the nineteenth and twentieth centuries. The corpus involves groups of mainly poetic works and authors that have a common denominator as regards schools, currents, culture and chronology.
Statistical information on a substantial corpus of representative Spanish texts is needed in order to determine the significance of data about individual authors or texts by means of comparison. This study describes the organization and analysis of a 150,000-word corpus of 30 well-known twentieth-century Spanish authors. Tables show the computational results of analyses involving sentences, segments, quotations, and word length.
This article summarizes the activities of the Istituto di Linguistica Computazionale. We discuss the Italian Multi-functional Lexical Databases; the projects focussing on linguistic analysis and generation; corpora in the MRF, textual databases and linguistic workstations; computer-assisted humanities teaching; and the various cooperative ventures, seminars and conferences offered by the Institute.
The Century of Prose Corpus is a historical corpus of British English of the period 1680–1780. It has been designed to provide a resource for students of the language of that era. The COPC is diachronic and may be considered a unit in what will eventually become a series of corpora providing access to the whole of the English language from the oldest specimens to the present. This article describes and explains the various features of the COPC.
This paper concerns the Charrette Project, a multimedia electronic archive of a medieval manuscript tradition. In this paper, we argue that the computer's strengths in manipulating complex and varied resources should be an important organizing principle in the conception and construction of electronic text projects. Specifically, we describe the elements of the Charrette archive, its architecture, and its potential for scholarly research and pedagogical applications.
This article is a detailed account of COMLEX Syntax, an on-line syntactic dictionary of English, developed by the Proteus Project at New York University under the auspices of the Linguistics Data Consortium. This lexicon was intended to be used for a variety of tasks in natural language processing by computer and as such has very detailed classes with a large number of syntactic features and complements for the major parts of speech and is, as far as possible, theory neutral. The dictionary was entered by hand with reference to hard copy dictionaries, an on-line concordance and native speakers‘intuition. Thus it is without prior encumbrances and can be used for both pure research and commercial purposes.
In this paper, we study the problem of adding a large number of new words into a Chinese thesaurus according to their definitions in a Chinese dictionary, while minimizing the effort of hand tagging. To deal with the problem, we first make use of a kind of supervised learning technique to learn a set of defining formats for each class in the thesaurus, which tries to characterize the regularities about the definitions of the words in the class. We then use traditional techniques in Graph theory to derive a minimal subset of the new words to be added into the thesaurus, which meets the following condition: if we add the new words in the subset into the thesaurus by hand, the other new words can be added into the thesaurus automatically by matching their definitions with the defining formats of each class in the thesaurus. The method uses little, if any, language-specific or thesaurus-specific knowledge, and can be applied to the thesauri of other languages.
This paper discusses the design of the EuroWordNet database, in which semantic databases like WordNet1.5 for several languages are combined via a so-called inter-lingual-index. In this database, language-independent data is shared whilst language-specific properties are maintained. A special interface has been developed to compare the semantic configurations across languages and to track down differences.
We discuss ways in which EuroWordNet (EWN) can be used in multilingual information retrieval activities, focusing on two approaches to Cross-Language Text Retrieval that use the EWN database as a large-scale multilingual semantic resource. The first approach indexes documents and queries in terms of the EuroWordNet Inter-Lingual-Index, thus turning term weighting and query/document matching into language-independent tasks. The second describes how the information in the EWN database could be integrated with a corpus-based technique, thus allowing retrieval of domain-specific terms that may not be present in our multilingual database. Our objective is to show the potential of EuroWordNet as a promising alternative to existing approaches to Cross-Language Text Retrieval.
This paper describes how the Euro WordNet project established a maximum level of consensus in the interpretation of relations, without loosing the possibility of encoding language-specific lexicalizations. Problematic cases arise due to the fact that each site re-used different resources and because the core vocabulary of the wordnets show complex properties. Many of these cases are discussed with respect to language internal and equivalence relations. Possible solutions are given in the form of additional criteria.
In this paper the linguistic design of the database under construction within the EuroWordNet project is described. This is mainly structured along the same lines as the Princeton WordNet, although some changes have been made to the WordNet overall design due to both theoretical and practical reasons. The most important reasons for such changes are the multilinguality of the EuroWordNet database and the fact that it is intended to be used in Language Engineering applications. Thus, i) some relations have been added to those identified in WordNet; ii) some labels have been identified which can be added to the relations in order to make their implications more explicit and precise; iii) some relations, already present in the WordNet design, have been modified in order to specify their role more clearly.
We give a brief outline of the design and contents of the English lexical database WordNet, which serves as a model for similarly conceived wordnets in several European languages. WordNet is a semantic network, in which the meanings of nouns, verbs, adjectives, and adverbs are represented in terms of their links to other (groups of) words via conceptual-semantic and lexical relations. Each part of speech is treated differently reflecting different semantic properties. We briefly discuss polysemy in WordNet, and focus on the case of meaning extensions in the verb lexicon. Finally, we outline the potential uses of WordNet not only for applications in natural language processing, but also for research in stylistic analyses in conjunction with a semantic concordance.
This paper describes two fundamental aspects in the process of building of the EuroWordNet database. In EuroWordNet we have chosen for a flexible design in which local wordnets are built relatively independently as language-specific structures, which are linked to an Inter-Lingual-Index (ILI). To ensure compatibility between the wordnets, a core set of common concepts has been defined that has to be covered by every language. Furthermore, these concepts have been classified via the ILI in terms of a Top Ontology of 63 fundamental semantic distinctions used in various semantic theories and paradigms. This paper first discusses the process leading to the definition of the set of Base Concepts, and the structure and the rationale of the Top Ontology.
This paper gives a global introduction to the aims and objectives of the EuroWordNet project, and it provides a general framework for the other papers in this volume. EuroWordNet is an EC project that develops a multilingual database with wordnets in several European languages, structured along the same lines as the Princeton WordNet. Each wordnet represents an autonomous structure of language-specific lexicalizations, which are interconnected via an Inter-Lingual-Index. The wordnets are built at different sites from existing resources, starting from a shared level of basic concepts and extended top-down. The results will be publicly available and will be tested in cross-language information retrieval applications.
INTEX is a linguistic development environment that includes large-coverage dictionaries and grammars, and parses texts of several million words in real time. INTEX has tools to create and maintain large-coverage lexical resources as well as morphological and syntactic grammars. Dictionaries and grammars are applied to texts in order to locate morphological, lexical and syntactic patterns, remove ambiguities, and tag simple and compound words. INTEX can build lemmatized concordances and indices of large texts with respect to all types of Finite State patterns. INTEX is used as a corpus processor, to analyze literary, journalistic and technical texts. I describe here the subset of tools used to perform advanced search requests on large texts.
We report on a project to annotate biblical texts in order to create an aligned multilingual Bible corpus for linguistic research, particularly computational linguistics, including automatically creating and evaluating translation lexicons and semantically tagged texts. The output of this project will enable researchers to take advantage of parallel translations across a wider number of languages than previously available, providing, with relatively little effort, a corpus that contains careful translations and reliable alignment at the near-sentence level. We discuss the nature of the text, our annotation process, preliminary and planned uses for the corpus, and relevant aspects of the Corpus Encoding Standard (CES) with respect to this corpus. We also present a quantitative comparison with dictionary and corpus resources for modern-day English, confirming the relevance of this corpus for research on present day language.
This paper presents some aspects of the Silfide server, a system dedicated to the delivery of linguistic resources on the web. After presenting the main issues behind the design of such a system, we focus on the editorial choices related to the use of the Text Encoding Initiative to represent our textual documents. In particular, we focus on the accommodations we have had to carry with regards to the TEI header and address the trade-off between extensive enrichment and genericity of the primary data when one wants to precisely mark-up a given document content. As a whole, we show how essential the TEI has proven to be for a project such as ours both from a practical and conceptual point of view.
The CL Research Senseval system wasthe highest performing system among the ``All-words''systems, with an overall fine-grained score of 61.6percent for precision and 60.5 percent for recall on98 percent of the 8,448 texts on the revisedsubmission (up by almost 6 and 9 percent from thefirst). The results were achieved with an almostcomplete reliance on syntactic behavior, using (1) arobust and fast ATN-style parser producing parse treeswith annotations on nodes, (2) DIMAP dictionarycreation and maintenance software (after conversion ofthe Hector dictionary files) to hold dictionaryentries, and (3) a strategy for analyzing the parsetrees in concert with the dictionary data. Furtherconsiderable improvements are possible in the parser,exploitation of the Hector data (and representation ofdictionary entries), and the analysis strategy, stillwith syntactic and collocational data. The Sensevaldata (the dictionary entries and the corpora) providean excellent testbed for understanding the sources offailures and for evaluating changes in the CL Researchsystem.
We describe how to build a largecomprehensive, integrated Arabic lexicon byautomatic parsing of newspaper text. We havebuilt a parser system to read Arabic newspaperarticles, isolate the tokens from them, findthe part of speech, and the features for eachtoken. To achieve this goal we designed a setof algorithms, we generated several sets ofrules, and we developed a set of techniques,and a set of components to carry out thesetechniques. As each sentence is processed, newwords and features are added to the lexicon, sothat it grows continuously as the system runs.To test the system we have used 100 articles(80,444 words) from the Al-Raya newspaper.The system consists of several modules: thetokenizer module to isolate the tokens, the type findersystem to find the part of speech of eachtoken, the proper noun phrase parser module tomark the proper nouns and to discover someinformation about them and the feature findermodule to find the features of the words.
As language data and associatedtechnologies proliferate and as the languageresources community expands, it is becomingincreasingly difficult to locate and reuse existingresources. Are there any lexical resources forsuch-and-such a language? What tool workswith transcripts in this particular format?What is a good format to use for linguisticdata of this type? Questions like these dominate manymailing lists, since web search engines are anunreliable way to find language resources. Thispaper reports on a new digital infrastructurefor discovering language resources beingdeveloped by the Open Language Archives Community(OLAC). At the core of OLAC is its metadataformat, which is designed to facilitatedescription and discovery of all kinds oflanguage resources, including data, tools, oradvice. The paper describes OLAC metadata, itsrelationship to Dublin Core metadata, and itsdissemination using the metadata harvesting protocol of the Open Archives Initiative.