Abstract The essential first step in the development of any quantitative method is identifying something to measure. If we want to use a quantitative technique to establish whether languages are likely to be related, or whether they fall into the same subgroup, or how similar two languages are relative to a third language, we need to decide what we are going to count; and to ensure that we are comparing like with like, the something we are counting has to be the same across all the languages we are comparing. Indeed, to make the method as flexible as possible we need that something to be the same in all the languages we might ever want to compare. In comparative linguistics this is a tall order. Our first thought might be to turn to sociolinguistics, where quantitative methods have enjoyed such success, and adopt the strategies developed there. However, sociolinguistic studies tend to operate within a single speech community (Patrick 2002), where speakers share the same norms of behaviour and attitude, and the same variable elements of linguistic structure. It is possible, therefore, to isolate a set of variables and to study the circumstances, both linguistic and non-linguistic, under which the different variants emerge. For example, we might consider a phonological variable (t), with variants [t], used mainly by women, and glottal stop [?], mainly used by men. We might find a syntactic variable (negation), with multiple negation (I didn’t do nothing ) favoured by lower-class speakers, and single negative markers (I didn’t do anything) the dominant variant for middle-class speakers; or a lexical variable, where the meaning (be sick) might be expressed by vomit for older speakers and throw up for younger ones.