The FUL (featurally underspecified lexicon) model of automatic speech recognition is based on the representation of words in the lexicon with underspecified distinctive features. The speech signal is converted from the waveform into an online spectral representation made up of formants and a few parameters describing the overall spectral shape. These LPC and spectral parameters are converted into distinctive phonological features which, in turn, are compared with all entries in the lexicon. No classification into segments, syllables, or spectral templates is used for the selection of words from the lexicon. Comparison of signal features with those stored in the lexicon uses a ternary system of matching, nomismatching, and mismatching features. Matching features increase the scoring for potential word candidates, no-mismatching features do not exclude candidates and only mismatching features lead to the rejection of word candidates. The word candidates are expanded to include word hypotheses, even without further acoustic evidence, and are used in the phonological and syntactic parsing that operates in parallel with the acoustic front-end. 1. THEORY The speech signal of the same phonetic segment varies across dialects and speakers — within speakers in certain segmental and prosodic contexts, and even for the same speaker and context with repetition, speaking rate, emotional state, microphones, etc. Not surprisingly, speech recognition with simple spectral template matching has failed consistently. Any variation in the signal leads to variation of the spectra that are compared to the stored templates. Only statistical approaches like Hidden Markov Models based on large training sets have led to acceptable results, but are still speaker and transmission-line dependent or operate only with a restricted vocabulary, syntax, and semantics. Human listeners seem to be unconcerned by adverse acoustical conditions and are able to resolve a wide range of variations like assimilations and deletions with apparent ease. Ambiguities in the signal, whether they come from random noise or whether they are linguistic in nature, like cliticizations of words or assimilations are the norm rather than the exception in natural language. Human listeners, however, appear not to be worried by adverse acoustic conditions and indeed, handle “variations” in the signal with ease. Language comprehension experiments [1, 2] have shown that listeners extract certain acoustic characteristics reliably but do not match acoustic details with the lexicon. Rather, the experimental results are best explained with the assumption that lexical access involves mapping the acoustic signal to an underspecified featural representation. For example, the assimilation of a coronal sound (e.g. /n/) to a following labial place of articulation (like [b] in “Where could Mr. Bean be?”) often results in the production of a labial (i.e. “Bea[m] be”). The reverse is not true, that is, a labial sound does not assimilate to a coronal place of articulation (i.e., “la[m]e duck” does not become “la[n]e duck”). Simple articulatory mechanics cannot account for such behaviour because an articulatory assimilation would operate in both directions. An explanation can be given by assuming that coronal sounds are underspecified for place, whereas labial and dorsals are not: the labial place of articulation spreads to the preceding coronal sound (if the language has regressive assimilation) because that sound is not specified for place. On the other hand, the specification of a labial place prevents the place features of an adjacent sound from overriding this information. Consequently, coronal sounds can become labial (or dorsal), but labial (or dorsal) sound cannot change their place. This explanation is straightforward for speech production, but what about speech perception? How can a realisation of “gree[m]” in a labial context (like “bag”) or “gree[N]” in a dorsal context (like “grass”) lead to the access of the word “green” in the lexicon? Normally, “gree[m]” and “gree[N]” are nonwords in English. And, at the same time, how should a mechanism be constructed to allow the activation of the word “bean” as well as “beam” if the acoustic input is “bea[m]”, when “bean” is a word of the language? Human listeners handle these asymmetries (and many other assimilatory effects) within and across words without noticing it, as reaction-time experiments have shown [4]. The solution to these seemingly contradictory requirements can be obtained (i) by assuming an underspecified representation in the lexicon, where certain features (like the place feature [coronal]) are not stored in the lexicon (in speech production, segments with unspecified place are generated with the feature coronal by default) and (ii) by postulating a ternary matching logic in the signal-to-lexical mapping. page 715 ICPhS99 San Francisco candidate 1 candidate 2 • • • candidate n 1