Most large language models are trained on linguistic input alone, yet humans\nappear to ground their understanding of words in sensorimotor experience. A\nnatural solution is to augment LM representations with human judgments of a\nword's sensorimotor associations (e.g., the Lancaster Sensorimotor Norms), but\nthis raises another challenge: most words are ambiguous, and judgments of words\nin isolation fail to account for this multiplicity of meaning (e.g., "wooden\ntable" vs. "data table"). We attempted to address this problem by building a\nnew lexical resource of contextualized sensorimotor judgments for 112 English\nwords, each rated in four different contexts (448 sentences total). We show\nthat these ratings encode overlapping but distinct information from the\nLancaster Sensorimotor Norms, and that they also predict other measures of\ninterest (e.g., relatedness), above and beyond measures derived from BERT.\nBeyond shedding light on theoretical questions, we suggest that these ratings\ncould be of use as a "challenge set" for researchers building grounded language\nmodels.\n
No citing papers are currently in WordNorms