Predicting Word-Naming and Lexical Decision Times from a Semantic Space Model Brendan T. Johns (johns4@indiana.edu) Department of Psychological and Brain Sciences, 1101 E. Tenth St. Bloomington, In 47405 USA Michael N. Jones (jonesmn@indiana.edu) Department of Psychological and Brain Sciences, 1101 E. Tenth St. Bloomington, In 47405 USA Abstract organization of semantic memory. In a lexical decision task, a letter string is presented and the participant provides a speeded response of whether the string is a word or not. In a naming task, the participant’s task is to name the presented word aloud as quickly as possible. Both measures produce an index of a word’s identification latency. Orthographic and phonological factors are certainly large components of both LDT and NT, but semantics plays a significant role as well, and co-occurrence models have yet to be extended to predicting reaction time variance for these single-word identification tasks. Modeling of retrieval times is usually done by looking for the best environmental correlates of LDT and NT (Adelman & Brown, 2008). Some of the most influential models of retrieval times are based upon word frequency. Word frequency (WF) has been used to drive many different types of models, including serial-searched rank frequency models (Murray & Forster, 2004), threshold activation models (Coltheart, et al., 2001), and connectionist models (Seidenberg & McClelland, 1989). However, recent evidence suggests that word frequency may not drive retrieval times but, rather, the causal factor is a word’s contextual diversity (Adelman, Brown, & Quesada, 2006; Adelman & Brown, 2008). Contextual diversity (CD) is the number of different contexts that a word appears in, and is based on the rational analysis of memory (Anderson & Milson, 1989), particularly the principle of likely need (PLN). PLN states that the more unique contexts a word appears in, the more likely the word will be needed in any future context. Hence, a word with a high CD should be faster to retrieve under this principle. A word’s CD value is typically computed by simply counting the number of different documents in which it appears across a text corpus. This measure has been shown to be a better predictor of LDT and NT than WF (Adelman, et al., 2006). However, operationalizing CD as the number of documents in which a word occurs may not be a fair instantiation of PLN. A word that appears in many documents may have a high WF, but it should have a low CD if those documents are highly redundant, as is the case with words that belong to a popular discourse topic for which many documents exist. It is the number of different contexts and the uniqueness of contexts that determines a word’s likely need. This calls for a measure of CD that considers the semantic uniqueness of documents that a word appears in. Based on PLN, it is reasonable to assume that if a word appears in a context it has never before occurred in, We propose a method to derive predictions for single-word retrieval times from a semantic space model trained on text corpora. In Experiment 1 we present a large corpus analysis demonstrating that it is the number of unique semantic contexts a word appears in across language, rather than simply the number of contexts or the frequency of the word, that is the most salient predictor of lexical decision and naming times. In Experiment 2, we develop a co-occurrence learning model that weights new contextual uses of a word based on fit to what currently exists in the word’s memory representation, and demonstrate this model’s superiority in fitting the human data compared to models built using information about the word’s frequency or number of contexts. Finally, in Experiment 3 we find that building lexical representations using semantic distinctiveness naturally produces a better-organized semantic space to make predictions for semantic similarity between words. Keywords: Co-occurrence model; Lexical-decision; LSA; Contextual distinctiveness Introduction The last decade has seen remarkable progress with co- occurrence models of lexical semantics (e.g., Lund & Burgess, 1996; Landauer & Dumais, 1997). These models learn semantic representations for words by observing lexical co-occurrence patterns across a large text corpus, typically representing the words in a high-dimensional semantic space. This approach provides both an account of the semantic representation for words and an account of the learning mechanisms humans use to build and organize semantic memory. Co-occurrence models have seen considerable success at accounting for data in a wide variety of semantic tasks, including TOEFL synonyms (Landauer & Dumais, 1997), semantic similarity ratings and exemplar categorization (Jones & Mewhort, 2007), and free association norms (Griffiths, Steyvers, & Tenenbaum, To date, all applications of co-occurrence models have been to semantic similarity between two words or two documents. The standard prediction of semantic similarity in these models is some measure of the angle between two vectors. However, co-occurrence models should, in theory, contain sufficient information in the magnitude of their representations to make predictions about single word retrieval as well. Lexical decision time (LDT) and word naming time (NT) are both important variables that offer insight into the