7.3.2 Predicting Phonetic Variation
For speech synthesis as well as recognition, we need to be able to represent the relation between the abstract category and its surface appearance, and predict the surface appearance from the abstract category and the context of the utterance. In early work in phonology, the relationship between a phoneme and its allophones was captured by writing a phonological rule. Here is the phonological rule for flapping in the traditional notation of Chomsky and Halle (1968):
$$ \left/\left\{\begin{array}{l}t\\ d\end{array}\right\}\right/\rightarrow[d\mathrm{x}]/\check{\mathrm{V}}\underline{\quad}\mathrm{V} $$
In this notation, the surface allophone appears to the right of the arrow, and the phonetic environment is indicated by the symbols surrounding the underbar (\_\). Simple rules like these are used in both speech recognition and synthesis when we want to generate many pronunciations for a word; in speech recognition this is often used as a first step toward picking the most likely single pronunciation for a word (see Sec. \ref{sec_sec}).
In general, however, there are two reasons why these simple 'Chomsky-Halle'-type rules don't do well at telling us when a given surface variant is likely to be used. First, variation is a stochastic process; flapping sometimes occurs, and sometimes doesn't, even in the same environment. Second, many factors that are not related to the phonetic environment are important to this prediction task. Thus linguistic research and speech recognition/synthesis both rely on statistical tools to predict the surface form of a word by showing which factors cause, e.g., a particular /t/ to flap in a particular context.