10.5.3 Pronunciation Modeling: Variation due to Genre
We said at the beginning of the chapter that recognizing conversational speech is harder for ASR systems than recognizing read speech. What are the causes of this difference? Is it the difference in vocabulary? Grammar? Something about the speaker themselves? Perhaps it's a fact about the microphones or telephone used in the experiment.
None of these seems to be the cause. In a well-known experiment, Weintraub et al. (1996) compared ASR performance on natural conversational speech versus performance on read speech, controlling for the influence of possible causal factors. Pairs of subjects in the lab had spontaneous conversations on the telephone. Weintraub et al. (1996) then hand-transcribed the conversations, and invited the participants back into the lab to read their own transcripts to each other over the same phone lines as if they were dictating. Both the natural and read conversations were recorded. Now Weintraub et al. (1996) had two speech corpora from identical transcripts; one original natural conversation, and one read speech. In both cases the speaker, the actual words, and the microphone were identical; the only difference was the naturalness or fluency of the speech. They found that read speech was much easier (WER=29%) than conversational speech (WER=53%). Since the speakers, words, and channel were controlled for, this difference must be modelable somewhere in the acoustic model or pronunciation lexicon.
Saraclar et al. (2000) tested the hypothesis that this difficulty with conversational speech was due to changed pronunciations, i.e., to a mismatch between the phone strings in the lexicon and what people actually said. Recall from Ch. 7 that conversational corpora like Switchboard contain many different pronunciations for words, (such as 12 different pronunciations for because and hundreds for the). Saraclar et al. (2000) showed in an oracle experiment that if a Switchboard recognizer is told which pronunciations to use for each word, the word error rate drops from 47% to 27%.
If knowing which pronunciation to use improves accuracy, could we improve recognition by simply adding more pronunciations for each word to the lexicon?
Alas, it turns out that adding multiple pronunciations doesn't work well, even if the list of pronunciation is represented as an efficient pronunciation HMM (Cohen, 1989). Adding extra pronunciations adds more confusability; if a common pronunciation of the word "of" is the single vowel [ax], it is now very confusable with the word "a". Another problem with multiple pronunciations is the use of Viterbi decoding. Recall our discussion on 2 that since the Viterbi decoder finds the best phone string, rather than the best word string, it biases against words with many pronunciations. Finally, using multiple pronunciations to model coarticulatory effects may be unnecessary because CD phones (triphones) are already quite good at modeling the contextual effects in phones due to neighboring phones, like the flapping and vowel-reduction handled by Fig. ?? (Jurafsky et al., 2001).
Instead, most current LVCSR systems use a very small number of pronunciations per word. What is commonly done is to start with a multiple pronunciation lexicon, where the pronunciations are found in dictionaries or are generated via phonological rules of the type described in Ch. 7. A forced Viterbi phone alignment is then run of the training set, using this dictionary. The result of the alignment is a phonetic transcription of the training corpus, showing which pronunciation was used, and the frequency of each pronunciation. We can then collapse similar pronunciations (for example if two pronunciations differ only in a single phone substitution we chose the more frequent pronunciation). We then chose the maximum likelihood pronunciation for each word. For frequent words which have multiple high-frequency pronunciations, some systems chose multiple pronunciations, and annotate the dictionary with the probability of these pronunciations; the probabilities are used in computing the acoustic likelihood (Cohen, 1989; Hain et al., 2001; Hain, 2002).
Finding a better method to deal with pronunciation variation remains an unsolved research problem. One promising avenue is to focus on non-phonetic factors that affect pronunciation. For example words which are highly predictable, or at the beginning or end of intonation phrases, or are followed by disfluencies, are pronounced very differently (Jurafsky et al., 1998; Fosler-Lussier and Morgan, 1999; Bell et al., 2003). Fosler-Lussier (1999) shows an improvement in word error rate by using these sorts of factors to predict which pronunciation to use. Another exciting line of research in pronunciation modeling uses a dynamic Bayesian network to model the complex overlap in articulators that produces phonetic reduction (Livescu and Glass, 2004b, 2004a).
Another important issue in pronunciation modeling is dealing with unseen words. In web-based applications such as telephone-based interfaces to the Web, the recognizer lexicon must be automatically augmented with pronunciations for the millions of unseen words, particularly names, that occur on the Web. Grapheme-to-phoneme techniques like those described in Sec. ?? are used to solve this problem.
10.6 METADATA: BOUNDARIES, PUNCTUATION, AND DISFLUENCIES
The output of the speech recognition process as we have described it so far is just a string of raw words. Consider the following sample gold-standard transcript (i.e., assuming perfect word recognition) of part of a dialogue (Jones et al., 2003):
yeah actually um i belong to a gym down here a gold's gym uh-huh and uh exercise i try to exercise five days a week um and i usually do that uh what type of exercising do you do in the gym
Compare the difficult transcript above with the following much clearer version:
A: Yeah I belong to a gym down here. Gold's Gym. And I try to exercise five days a week. And I usually do that.
B: What type of exercising do you do in the gym?
The raw transcript is not divided up among speakers, there is no punctuation or capitalization, and disfluencies are scattered among the words. A number of studies have shown that such raw transcripts are harder for people to read Jones et al. (2003, 2005) and that adding, for example, commas back into the transcript improve the accuracy of information extraction algorithms on the transcribed text (Makhoul et al., 2005; Hillard et al., 2006). Post-processing ASR output involves tasks including the following:
diarization: Many speech tasks have multiple speakers, such as telephone conversations, business meetings, and news reports (with multiple broadcasters). Diarization is the task of breaking up a speech file by speaker assigning parts of the transcript to the relevant speakers, like the A: and B: labels above.
sentence boundary detection: We discussed the task of breaking speech into sentences (sentence segmentation) in Ch. 3 and Ch. 8. But for those tasks we already add punctuation like periods to help us; from speech we don't already have punctuation, just words. Sentence segmentation from speech has the added difficulty that the transcribed words will be errorful, but has the advantage that prosodic features like pauses and sentence-final intonation can be used as cues.
truecasing: Words in a clean transcript need to have sentence-initial words starting with an upper-case letter, acronyms all in capitals, and so on. Truecasing is the task of assigning the correct case for a word, and is often addressed as a HMM classification task like part-of-speech tagging, with hidden states like ALL-LOWER CASE, UPPER-CASE-INITIAL, all-caps, and so on.
punctuation detection: In addition to segmenting sentences, we need to choose sentence-final punctuation (period, question mark, exclamation mark), and insert commas and quotation marks and so on.
disfluency detection: Disfluencies can be removed from a transcript for readability, or at least marked off with commas or font changes. Since standard recognizers don't actually include disfluencies (like word fragments) in their transcripts, disfluency detection algorithms can also play an important role in avoiding the misrecognized words that may result.
Marking these features (punctuation, boundaries, diarization) in the text output is often called metadata or sometimes rich transcription. Let's look at a couple of these tasks in slightly more detail.
Sentence segmentation can be modeled as a binary classification task, in which each boundary between two words is judged as a sentence boundary or as sentence-internal. Such classifiers can use similar features to the sentence segmentation discussed in Sec. \ref{sec:text-segmentation}. Such classifiers can use similar features to the sentence segmentation discussed in Sec. \ref{sec:text-segmentation}. Such classifiers can use similar features to the sentence segmentation discussed in Sec. \ref{sec:text-segmentation}.
Sentence segmentation can be modeled as a binary classification task, in which each boundary between two words is judged as a sentence boundary or as sentence-internal. Such classifiers can use similar features to the sentence segmentation discussed in Sec. \ref{sec:text-segmentation}. Such classifiers can use similar features to the sentence segmentation discussed in Sec. \ref{sec:text-segmentation}.
Sentence segmentation can be modeled as a binary classification task, in which each boundary between two words is judged as a sentence boundary or as sentence-internal. Such classifiers can use similar features to the sentence segmentation discussed in Sec. \ref{sec:text-segmentation}. Such classifiers can use similar features to the sentence segmentation discussed in Sec. \ref{sec:text-segmentation}.
Fig. 10.17 shows the candidate boundary locations in a sample sentence. Commonly extracted features include:
pause features: duration of the interword pause at the candidate boundary.

duration features: durations of the phone and rime (nucleus plus coda) preceding the candidate boundary. Since some phones are inherently longer than others, each phone is normalized to the mean duration for that phone.
F0 features: the change in pitch across the boundary; sentence boundaries often have pitch reset (an abrupt change in pitch), while non-boundaries are more likely to have continuous pitch across the boundary. Another useful F0 feature is the pitch range of the preboundary word; sentences often end with a final fall (Sec. ??) which is close to the speaker's F0 baseline.
For punctuation detection, similar features are used as for sentence boundary detection, but with multiple hidden classes (comma, sentence-final question mark, quotation mark, no punctuation). Instead of just two.
For both of these tasks, instead of a simple binary classifier, sequence information can be incorporated by modeling sentence segmentation as an HMM in which the hidden states correspond to sentence boundary or non-boundary decisions. We will describe methods for combining prosodic and lexical features in more detail when we introduce dialogue act detection in Sec. ??.
Recall from Sec. ?? that disfluencies or repair in conversation include phenomena like the following:
| Disfluency type | Example |
| fillers (or filled pauses): | But, uh, that was absurd |
| word fragments | A guy went to a d-, a landfill |
| repetitions: | it was just a change of, change of location |
| restarts | it's - I find it very strange |
The ATIS sentence in Fig. 10.18 shows examples of a restart and the filler uh, showing the
Detection methods for disfluencies are very similar to detecting sentence boundaries; a classifier is trained to make a decision at each word boundary, using both text and prosodic features. HMM and CRF classifiers are commonly used, and features are quite similar to the features for boundary detection, including neighboring words and part-of-speech tags, the duration of pauses at the word boundary, the duration of the word and phones preceding the boundary, the difference in pitch values across the boundary, and so on.
For detecting fragments, features for detecting voice quality are used (Liu, 2004), such as jitter, a measure of perturbation in the pitch period (Rosenberg, 1971), spectral
| Interruption Point | |
| Does American Airlines offer any one-way flights | [uh] one-way fares for 160 dollars? |
| Reparandum | Repair |
| Editing Phase | |
| Figure 10.18 Repeated from Fig. ??An example of a disfluency (after Shriberg (1994); terminology is from Levelt (1983)). | |

tilt, the slope of the spectrum, (see Sec. ??), and open quotient, the percentage of the glottal cycle in which the vocal folds are open (Fant, 1997).
10.7 SPEECH RECOGNITION BY HUMANS
Speech recognition in humans shares some features with ASR algorithms. We mentioned above that signal processing algorithms like PLP analysis (Hermansky, 1990) were in fact inspired by properties of the human auditory system. In addition, three properties of human lexical access (the process of retrieving a word from the mental lexicon) are also true of ASR models: frequency, parallelism, and cue-based processing. For example, as in ASR with its N-gram language models, human lexical access is sensitive to word frequency. High-frequency spoken words are accessed faster or with less information than low-frequency words. They are successfully recognized in noisier environments than low frequency words, or when only parts of the words are presented (Howes, 1957; Grosjean, 1980; Tyler, 1984, inter alia). Like ASR models, human lexical access is parallel: multiple words are active at the same time (Marslen-Wilson and Welsh, 1978; Salasoo and Pisoni, 1985, inter alia).
Humans are of course much better at speech recognition than machines; current machines are roughly about five times worse than humans on clean speech, and the gap seems to increase with noisy speech.
Finally, human speech perception is cue based: speech input is interpreted by integrating cues at many different levels. Human phone perception combines acoustic cues, such as formant structure or the exact timing of voicing, (Oden and Massaro, 1978; Miller, 1994) visual cues, such as lip movement (McGurk and Macdonald, 1976; Massaro and Cohen, 1983; Massaro, 1998) and lexical cues such as the identity of the word in which the phone is placed (Warren, 1970; Samuel, 1981; Connie and Clifton, 1987; Connine, 1990). For example, in what is often called the phoneme restoration effect, Warren (1970) took a speech sample and replaced one phone (e.g. the [s] in legislature) with a cough. Warren found that subjects listening to the resulting tape typically heard the entire word legislature including the [s], and perceived the cough as background. In the McGurk effect, (McGurk and Macdonald, 1976) showed that visual input can interfere with phone perception, causing us to perceive a completely different phone. They showed subjects a video of someone saying
ing the syllable ga in which the audio signal was dubbed instead with someone saying the syllable ba. Subjects reported hearing something like da instead. It is definitely worth trying this out yourself from video demos on the web; see for example http://www.haskins.yale.edu/featured/heads/mcgurk.html. Other cues in human speech perception include semantic word association (words are accessed more quickly if a semantically related word has been heard recently) and repetition priming (words are accessed more quickly if they themselves have just been heard). The intuitions of both these results are incorporated into recent language models discussed in Ch. 4, such as the cache model of Kuhn and De Mori (1990), which models repetition priming, or the trigger model of Rosenfeld (1996) and the LSA models of Coccaro and Jurafsky (1998) and Bellegarda (1999), which model word association. In a fascinating reminder that good ideas are never discovered only once, Cole and Rudnicky (1983) point out that many of these insights about context effects on word and phone processing were actually discovered by William Bagley (1901). Bagley achieved his results, including an early version of the phoneme restoration effect, by recording speech on Edison phonograph cylinders, modifying it, and presenting it to subjects. Bagley's results were forgotten and only rediscovered much later. $ ^{2} $
One difference between current ASR models and human speech recognition is the time-course of the model. It is important for the performance of the ASR algorithm that the decoding search optimizes over the entire utterance. This means that the best sentence hypothesis is returned by a decoder at the end of the sentence may be very different than the current-best hypothesis, halfway into the sentence. By contrast, there is extensive evidence that human processing is on-line: people incrementally segment and utterance into words and assign it an interpretation as they hear it. For example, Marslen-Wilson (1973) studied close shadowers: people who are able to shadow (repeat back) a passage as they hear it with lags as short as 250 ms. Marslen-Wilson found that when these shadowers made errors, they were syntactically and semantically appropriate with the context, indicating that word segmentation, parsing, and interpretation took place within these 250 ms. Cole (1973) and Cole and Jakimik (1980) found similar effects in their work on the detection of mispronunciations. These results have led psychological models of human speech perception (such as the Cohort model (Marslen-Wilson and Welsh, 1978) and the computational TRACE model (McClelland and Elman, 1986)) to focus on the time-course of word selection and segmentation. The TRACE model, for example, is a connectionist interactive-activation model, based on independent computational units organized into three levels: feature, phoneme, and word. Each unit represents a hypothesis about its presence in the input. Units are activated in parallel by the input, and activation flows between units; connections between units on different levels are excitatory, while connections between units on single level are inhibitory. Thus the activation of a word slightly inhibits all other words.
We have focused on the similarities between human and machine speech recognition; there are also many differences. In particular, many other cues have been shown to play a role in human speech recognition but have yet to be successfully integrated into ASR. The most important class of these missing cues is prosody. To give only one example, Cutler and Norris (1988), Cutler and Carter (1987) note that most mul
tisyllabic English word tokens have stress on the initial syllable, suggesting in their metrical segmentation strategy (MSS) that stress should be used as a cue for word segmentation. Another difference is that human lexical access exhibits neighborhood effects (the neighborhood of a word is the set of words which closely resemble it). Words with large frequency-weighted neighborhoods are accessed slower than words with less neighbors (Luce et al., 1990). Current models of ASR don't generally focus on this word-level competition.
10.8 SUMMARY
• We introduced two advanced decoding algorithms: The multipass (N-best or lattice) decoding algorithm, and stack or A $ ^{*} $ decoding.
- Advanced acoustic models are based on context-dependent triphones rather than phones. Because the complete set of triphones would be too large, we use a smaller number of automatically clustered triphones instead.
Acoustic models can be adapted to new speakers.
- Pronunciation variation is a source of errors in human-human speech recognition, but one that is not successfully handled by current technology.
BIBLIOGRAPHICAL AND HISTORICAL NOTES
See the previous chapter for most of the relevant speech recognition history. Note that although stack decoding is equivalent to the A* search developed in artificial intelligence, the stack decoding algorithm was developed independently in the information theory literature and the link with AI best-first search was noticed only later (Jelinek, 1976). Useful references on vocal tract length normalization include (Cohen et al., 1995; Wegmann et al., 1996; Eide and Gish, 1996; Lee and Rose, 1996; Welling et al., 2002; Kim et al., 2004).
There are many new directions in current speech recognition research involving alternatives to the HMM model. For example, there are new architectures based on graphical models (dynamic bayes nets, factorial HMMs, etc) (Zweig, 1998; Bilmes, 2003; Livescu et al., 2003; Bilmes and Bartels, 2005; Frankel et al., 2007). There are attempts to replace the frame-based HMM acoustic model (that makes a decision about each frame) with segment-based recognizers that attempt to detect variable-length segments (phones) (Digilakis, 1992; Ostendorf et al., 1996; Glass, 2003). Landmark-based recognizers and articulatory phonology-based recognizers focus on the use of distinctive features, defined acoustically or articulatorily (respectively) (Niyogi et al., 1998; Livescu, 2005; Hasegawa-Johnson and et al., 2005; Juneja and Espy-Wilson, 2003).
See Shriberg (2005) for an overview of metadata research. Shriberg (2002) and Nakatani and Hirschberg (1994) are computationally-focused corpus studies of the acoustic and lexical properties of disfluencies. Early papers on sentence segmentation
tion from speech include Wang and Hirschberg (1992), Ostendorf and Ross (1997) See Shriberg et al. (2000), Liu et al. (2006a) for recent work on sentence segmentation, Kim and Woodland (2001), Hillard et al. (2006) on punctuation detection, Nakatani and Hirschberg (1994), Honal and Schultz (2003, 2005), Lease et al. (2006), and a number of papers that jointly address multiple metadata extraction tasks (Heeman and Allen, 1999; Liu et al., 2005, 2006b).
EXERCISES
10.1 Implement the Stack decoding algorithm of Fig. 10.7 on page 9. Pick a very simple $ h^{*} $ function like an estimate of the number of words remaining in the sentence.
10.2 Modify the forward algorithm of Fig. ?? from Ch. 9 to use the tree-structured lexicon of Fig. 10.10 on page 12.
10.3 Many ASR systems, including the Sonic and HTK systems, use a different algorithm for Viterbi called the token-passing Viterbi algorithm (Young et al., 1989). Read this paper and implement this algorithm.
Atal, B. S. (1974). Effectiveness of linear prediction characteristics of the speech wave for automatic speaker identification and verification. The Journal of the Acoustical Society of America, 55(6), 1304–1312.
Aubert, X. and Ney, H. (1995). Large vocabulary continuous speech recognition using word graphs. In IEEE ICASSP, Vol. 1, pp. 49–52.
Austin, S., Schwartz, R., and Placeway, P. (1991). The forward-backward search algorithm. In IEEE ICASSP-91, Vol. 1, pp. 697–700.
Bagley, W. C. (1900–1901). The apperception of the spoken sentence: A study in the psychology of language. The American Journal of Psychology, 12, 80–130. †.
Bahl, L. R., Brown, P. F., de Souza, P. V., and Mercer, R. L. (1986). Maximum mutual information estimation of hidden Markov model parameters for speech recognition. In IEEE ICASSP-86, Tokyo, pp. 49–52.
Bahl, L. R., de Souza, P. V., Gopalakrishnan, P. S., Nahamoo, D., and Picheny, M. A. (1992). A fast match for continuous speech recognition using allophonic models. In IEEE ICASSP-92, San Francisco, CA, pp. I.17–20.
Bell, A., Jurafsky, D., Fosler-Lussier, E., Girand, C., Gregory, M. L., and Gildea, D. (2003). Effects of disfluencies, predictability, and utterance position on word form variation in English conversation. Journal of the Acoustical Society of America, 113(2), 1001–1024.
Bellegarda, J. R. (1999). Speech recognition experiments using multi-span statistical language models. In IEEE ICASSP-99, pp. 717–720.
Bilmes, J. (2003). Buried Markov Models: A graphical-modeling approach to automatic speech recognition. Computer Speech and Language, 17(2-3).
Bilmes, J. and Bartels, C. (2005). Graphical model architectures for speech recognition. IEEE Signal Processing Magazine, 22(5), 89–100.
Bourlard, H. and Morgan, N. (1994). Connectionist Speech Recognition: A Hybrid Approach. Kluwer Press.
Chou, W., Lee, C. H., and Juang, B. H. (1993). Minimum error rate training based on n-best string models. In IEEE ICASSP-93, pp. 2.652–655.
Coccaro, N. and Jurafsky, D. (1998). Towards better integration of semantic predictors in statistical language modeling. In ICSLP-98, Sydney, Vol. 6, pp. 2403–2406.
Cohen, J., Kamm, T., and Andreou, A. (1995). Vocal tract normalization in speech recognition: compensating for systematic systematic speaker variability. Journal of the Acoustical Society of America, 97(5), 3246–3247.
Cohen, M. H. (1989). Phonological Structures for Speech Recognition. Ph.D. thesis, University of California, Berkeley.
Cole, R. A. (1973). Listening for mispronunciations: A measure of what we hear during speech. Perception and Psychophysics, 13, 153–156.
Cole, R. A. and Jakimik, J. (1980). A model of speech perception. In Cole, R. A. (Ed.), Perception and Production of Fluent Speech, pp. 133–163. Lawrence Erlbaum.
Cole, R. A. and Rudnicky, A. I. (1983). What's new in speech perception? The research and ideas of William Chandler Bagley. Psychological Review, 90(1), 94–101.
Connine, C. M. (1990). Effects of sentence context and lexical knowledge in speech processing. In Altmann, G. T. M. (Ed.), Cognitive Models of Speech Processing, pp. 281–294. MIT Press.
Connine, C. M. and Clifton, C. (1987). Interactive use of lexical information in speech perception. Journal of Experimental Psychology: Human Perception and Performance, 13, 291–299.
Cutler, A. and Carter, D. M. (1987). The predominance of strong initial syllables in the English vocabulary. Computer Speech and Language, 2, 133–142.
Cutler, A. and Norris, D. (1988). The role of strong syllables in segmentation for lexical access. Journal of Experimental Psychology: Human Perception and Performance, 14, 113–121.
Deng, L., Lennig, M., Seitz, F., and Mermelstein, P. (1990). Large vocabulary word recognition using context-dependent allophonic hidden Markov models. Computer Speech and Language, 4, 345–357.
Digilakis, V. (1992). Segment-based stochastic models of spectral dynamics for continuous speech recognition. Ph.D. thesis, Boston University.
Doumpiotis, V., Tsakalidis, S., and Byrne, W. (2003a). Discriminative training for segmental minimum bayes-risk decoding. In IEEE ICASSP-03.
Doumpiotis, V., Tsakalidis, S., and Byrne, W. (2003b). Lattice segmentation and minimum bayes risk discriminative training. In EUROSPEECH-03.
Eide, E. M. and Gish, H. (1996). A parametric approach to vocal tract length normalization. In IEEE ICASSP-96, Atlanta, GA, pp. 346–348.
Evermann, G. and Woodland, P. C. (2000). Large vocabulary decoding and confidence estimation using word posterior probabilities. In IEEE ICASSP-00, Istanbul, Vol. III, pp. 1655–1658.
Fant, G. (1997). The voice source in connected speech. Speech Communication, 22(2-3), 125–139.
Fosler-Lussier, E. (1999). Multi-level decision trees for static and dynamic pronunciation models. In EUROSPEECH-99, Budapest.
Fosler-Lussier, E. and Morgan, N. (1999). Effects of speaking rate and word predictability on conversational pronunciations. Speech Communication, 29(2-4), 137–158.
Frankel, J., Wester, M., and King, S. (2007). Articulatory feature recognition using dynamic bayesian networks. Computer Speech and Language, 21(4), 620–640.
Glass, J. (2003). A probabilistic framework for segment-based speech recognition. Computer Speech and Language, 17(1–2), 137–152.
Grosjean, F. (1980). Spoken word recognition processes and the gating paradigm. Perception and Psychophysics, 28, 267–283.
Gupta, V., Lennig, M., and Mermelstein, P. (1988). Fast search strategy in a large vocabulary word recognizer. Journal of the Acoustical Society of America, 84(6), 2007–2017.
Hain, T. (2002). Implicit pronunciation modelling in asr. In Proceedings of ISCA Pronunciation Modeling Workshop.
Hain, T., Woodland, P. C., Evermann, G., and Povey, D. (2001). New features in the CU-HTK system for transcription of conversational telephone speech. In IEEE ICASSP-01, Salt Lake City, Utah.
Hasegawa-Johnson, M. and et al (2005). Landmark-based speech recognition: Report of the 2004 Johns Hopkins Summer Workshop. In IEEE ICASSP-05.
Heeman, P. A. and Allen, J. (1999). Speech repairs, intonational phrases and discourse markers: Modeling speakers' utterances in spoken dialog. Computational Linguistics, 25(4).
Hermansky, H. (1990). Perceptual linear predictive (PLP) analysis of speech. Journal of the Acoustical Society of America, 87(4), 1738–1752.
Hillard, D., Huang, Z., Ji, H., Grishman, R., Hakkani-Tür, D., Harper, M., Ostendorf, M., and Wang, W. (2006). Impact of automatic comma prediction on pos/name tagging of speech. In Proceedings of IEEE/ACL 06 Workshop on Spoken Language Technology, Aruba.
Honal, M. and Schultz, T. (2003). Correction of disfluencies in spontaneous speech using a noisy-channel approach. In EUROSPEECH-03.
Honal, M. and Schultz, T. (2005). Automatic disfluency removal on recognized spontaneous speech - rapid adaptation to speaker-dependent disfluencies. In IEEE ICASSP-05.
Howes, D. (1957). On the relation between the intelligibility and frequency of occurrence of English words. Journal of the Acoustical Society of America, 29, 296–305.
Huang, C., Chang, E., Zhou, J., and Lee, K.-F. (2000). Accent modeling based on pronunciation dictionary adaptation for large vocabulary mandarin speech recognition. In ICSLP-00, Beijing, China.
Jelinek, F. (1969). A fast sequential decoding algorithm using a stack. IBM Journal of Research and Development, 13, 675–685.
Jelinek, F. (1976). Continuous speech recognition by statistical methods. Proceedings of the IEEE. 64(4). 532–557.
Jelinek, F. (1997). Statistical Methods for Speech Recognition. MIT Press.
Jelinek, F., Mercer, R. L., and Bahl, L. R. (1975). Design of a linguistic statistical decoder for the recognition of continuous speech. IEEE Transactions on Information Theory, IT-21(3), 250–256.
Jones, D. A., Gibson, E., Shen, W., Granoien, N., Herzog, M., Reynolds, D., and Weinstein, C. (2005). Measuring human readability of machine generated text: Three case studies in
speech recognition and machine translation. In IEEE ICASSP-05, pp. 18–23.
Jones, D. A., Wolf, F., Gibson, E., Williams, E., Fedorenko, E., Reynolds, D. A., and Zissman, M. (2003). Measuring the readability of automatic speech-to-text transcripts. In EUROSPEECH-03, pp. 1585–1588.
Juneja, A. and Espy-Wilson, C. (2003). Speech segmentation using probabilistic phonetic feature hierarchy and support vector machines. In IJCNN 2003.
Junqua, J. C. (1993). The Lombard reflex and its role on human listeners and automatic speech recognizers. Journal of the Acoustical Society of America, 93(1), 510–524.
Jurafsky, D., Ward, W., Jianping, Z., Herold, K., Xiuyang, Y., and Sen, Z. (2001). What kind of pronunciation variation is hard for triphones to model?. In IEEE ICASSP-01, Salt Lake City, Utah, pp. I.577–580.
Jurafsky, D., Bell, A., Fosler-Lussier, E., Girand, C., and Raymond, W. D. (1998). Reduction of English function words in Switchboard. In ICSLP-98, Sydney, Vol. 7, pp. 3111–3114.
Kim, D., Gales, M., Hain, T., and Woodland, P. C. (2004). Using vtn for broadcast news transcription. In ICSLP-04, Jeju, South Korea.
Kim, J. and Woodland, P. (2001). The use of prosody in a combined system for punctuation generation and speech recognition. In EUROSPEECH-01, pp. 2757–2760.
Klovstad, J. W. and Mondshein, L. F. (1975). The CASPERS linguistic analysis system. IEEE Transactions on Acoustics, Speech, and Signal Processing, ASSP-23(1), 118–123.
Kuhn, R. and De Mori, R. (1990). A cache-based natural language model for speech recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 12(6), 570–583.
Kumar, S. and Byrne, W. (2002). Risk based lattice cutting for segmental minimum Bayes-risk decoding. In ICSLP-02, Denver, CO.
Lease, M., Johnson, M., and Charniak, E. (2006). Recognizing disfluencies in conversational speech. IEEE Transactions on Audio, Speech and Language Processing, 14(5), 1566–1573.
Lee, L. and Rose, R. C. (1996). Speaker normalisation using efficient frequency warping procedures. In ICASSP96, pp. 353–356.
Leggetter, C. J. and Woodland, P. C. (1995). Maximum likelihood linear regression for speaker adaptation of HMMs. Computer Speech and Language, 9(2), 171–186.
Levelt, W. J. M. (1983). Monitoring and self-repair in speech. Cognition, 14, 41–104.
Liu, Y., Chawla, N. V., Harper, M. P., Shriberg, E., and Stolcke, A. (2006a). A study in machine learning from imbalanced data for sentence boundary detection in speech. Computer Speech & Language, 20(4), 468–494.
Liu, Y., Shriberg, E., Stolcke, A., Hillard, D., Ostendorf, M., and Harper, M. (2006b). Enriching speech recognition with automatic detection of sentence boundaries and disfluencies. IEEE Transactions on Audio, Speech, and Language Processing, 14(5), 1526–1540.
Liu, Y., Shriberg, E., Stolcke, A., Peskin, B., Ang, J., Hillard, D., Ostendorf, M., Tomalin, M., Woodland, P. C., and Harper, M. P. (2005). Structural metadata research in the ears program. In IEEE ICASSP-05.
Liu, Y. (2004). Word fragment identification using acoustic-prosodic features in conversational speech. In HLT-NAACL-03 student research workshop, pp. 37–42.
Livescu, K., Glass, J., and Bilmes, J. (2003). Hidden feature modeling for speech recognition using dynamic bayesian networks. In EUROSPEECH-03.
Livescu, K. (2005). Feature-Based Pronunciation Modeling for Automatic Speech Recognition. Ph.D. thesis, Massachusetts Institute of Technology.
Livescu, K. and Glass, J. (2004a). Feature-based pronunciation modeling for speech recognition. In HLT-NAACL-04, Boston, MA.
Livescu, K. and Glass, J. (2004b). Feature-based pronunciation modeling with trainable asynchrony probabilities. In ICSLP-04, Jeju, South Korea.
Luce, P. A., Pisoni, D. B., and Goldfinger, S. D. (1990). Similarity neighborhoods of spoken words. In Altmann, G. T. M. (Ed.), Cognitive Models of Speech Processing, pp. 122–147. MIT Press.
Makhoul, J., Baron, A., Bulyko, I., Nguyen, L., Ramshaw, L., Stallard, D., Schwartz, R., and Xiang, B. (2005). The effects of speech recognition and punctuation on information extraction performance. In INTERSPEECH-05, Lisbon, Portugal, pp. 57–60.
Mangu, L., Brill, E., and Stolcke, A. (2000). Finding consensus in speech recognition: Word error minimization and other applications of confusion networks. Computer Speech and Language, 14(4), 373–400.
Marslen-Wilson, W. and Welsh, A. (1978). Processing interactions and lexical access during word recognition in continuous speech. Cognitive Psychology, 10, 29–63.
Marslen-Wilson, W. (1973). Linguistic structure and speech shadowing at very short latencies. Nature, 244, 522–523.
Massaro, D. W. (1998). Perceiving Talking Faces: From Speech Perception to a Behavioral Principle. MIT Press.
Massaro, D. W. and Cohen, M. M. (1983). Evaluation and integration of visual and auditory information in speech perception. Journal of Experimental Psychology: Human Perception and Performance, 9, 753–771.
McClelland, J. L. and Elman, J. L. (1986). Interactive processes in speech perception: The TRACE model. In McClelland, J. L., Rumelhart, D. E., and the PDP Research Group (Eds.), Parallel Distributed Processing Volume 2: Psychological and Biological Models, pp. 58–121. MIT Press.
McDermott, E. and Hazen, T. (2004). Minimum Classification Error training of landmark models for real-time continuous speech recognition. In IEEE ICASSP-04.
McGurk, H. and Macdonald, J. (1976). Hearing lips and seeing voices. Nature, 264, 746–748.
Miller, J. L. (1994). On the internal structure of phonetic categories: a progress report. Cognition, 50, 271–275.
Murveit, H., Butzberger, J. W., Digalakis, V. V., and Weintraub, M. (1993). Large-vocabulary dictation using SRI's decipher speech recognition system: Progressive-search techniques. In IEEE ICASSP-93, Vol. 2, pp. 319–322. IEEE.
Nadas, A. (1983). A decision theoretic formulation of a training problem in speech recognition and a comparison of training by unconditional versus conditional maximum likelihood. IEEE Transactions on Acoustics, Speech, and Signal Processing, 31(4), 814–817.
Nakatani, C. and Hirschberg, J. (1994). A corpus-based study of repair cues in spontaneous speech. Journal of the Acoustical Society of America, 95(3), 1603–1616.
Ney, H., Haeb-Umbach, R., Tran, B.-H., and Oerder, M. (1992). Improvements in beam search for 10000-word continuous speech recognition. In IEEE ICASSP-92, San Francisco, CA, pp. 1.9–12. IEEE.
Nguyen, L. and Schwartz, R. (1999). Single-tree method for grammar-directed search. In IEEE ICASSP-99, pp. 613–616. IEEE.
Nilsson, N. J. (1980). Principles of Artificial Intelligence. Morgan Kaufmann, Los Altos, CA.
Niyogi, P., Burges, C., and Ramesh, P. (1998). Distinctive feature detection using support vector machines. In IEEE ICASSP-98.
Normandin, Y. (1996). Maximum mutual information estimation of hidden Markov models. In Lee, C. H., Soong, F. K., and Paliwal, K. K. (Eds.), Automatic Speech and Speaker Recognition, pp. 57–82. Kluwer.
Odell, J. J. (1995). The Use of Context in Large Vocabulary Speech Recognition. Ph.D. thesis, Queen's College, University of Cambridge.
Oden, G. C. and Massaro, D. W. (1978). Integration of featureful information in speech perception. Psychological Review, 85, 172–191.
Ortmanns, S., Ney, H., and Aubert, X. (1997). A word graph algorithm for large vocabulary continuous speech recognition. Computer Speech and Language, 11, 43–72.
Ostendorf, M., Digilakis, V., and Kimball, O. (1996). From HMMs to segment models: A unified view of stochastic modeling for speech recognition. IEEE Transactions on Speech and Audio, 4(5), 360–378.
Ostendorf, M. and Ross, K. (1997). Multi-level recognition of intonation labels. In Sagisaka, Y., Campbell, N., and Higuchi, N. (Eds.), Computing Prosody: Computational Models for Processing Spontaneous Speech, chap. 19, pp. 291–308. Springer.
Paul, D. B. (1991). Algorithms for an optimal A* search and linearizing the search in the stack decoder. In IEEE ICASSP-91, Vol. 1, pp. 693–696. IEEE.
Pearl, J. (1984). Heuristics. Addison-Wesley, Reading, MA.
Ravishankar, M. K. (1996). Efficient Algorithms for Speech Recognition. Ph.D. thesis, School of Computer Science, Carnegie Mellon University, Pittsburgh. Available as CMU CS tech report CMU-CS-96-143.
Rosenberg, A. E. (1971). Effect of Glottal Pulse Shape on the Quality of Natural Vowels. The Journal of the Acoustical Society of America, 49, 583–590.
Rosenfeld, R. (1996). A maximum entropy approach to adaptive statistical language modeling. Computer Speech and Language, 10, 187–228.
Salasoo, A. and Pisoni, D. B. (1985). Interaction of knowledge sources in spoken word identification. Journal of Memory and Language, 24, 210–231.
Samuel, A. G. (1981). Phonemic restoration: Insights from a new methodology. Journal of Experimental Psychology: General, 110, 474–494.
Saraclar, M., Nock, H., and Khudanpur, S. (2000). Pronunciation modeling by sharing gaussian densities across phonetic models. Computer Speech and Language, 14(2), 137–160.
Schwartz, R. and Austin, S. (1991). A comparison of several approximate algorithms for finding multiple (N-BEST) sentence hypotheses. In icassp91, Toronto, Vol. 1, pp. 701–704. IEEE.
Schwartz, R. and Chow, Y.-L. (1990). The N-best algorithm: An efficient and exact procedure for finding the N most likely sentence hypotheses. In IEEE ICASSP-90, Vol. 1, pp. 81–84. IEEE.
Schwartz, R., Chow, Y.-L., Kimball, O., Roukos, S., Krasnwer, M., and Makhoul, J. (1985). Context-dependent modeling for acoustic-phonetic recognition of continuous speech. In IEEE ICASSP-85, Vol. 3, pp. 1205–1208. IEEE.
Shriberg, E. (2002). To ‘errrr’ is human: ecology and acoustics of speech disfluencies. Journal of the International Phonetic Association, 31(1), 153–169.
Shriberg, E. (2005). Spontaneous speech: How people really talk, and why engineers should care. In INTERSPEECH-05, Lisbon, Portugal.
Shriberg, E., Stolcke, A., Hakkani-Tür, D., and Tür, G. (2000). Prosody-based automatic segmentation of speech into sentences and topics. Speech Communication, 32(1-2), 127–154.
Shriberg, E. (1994). Preliminaries to a Theory of Speech Disfluencies. Ph.D. thesis, University of California, Berkeley, CA. (unpublished).
Soong, F. K. and Huang, E.-F. (1990). A tree-trellis based fast search for finding the n-best sentence hypotheses in continuous speech recognition. In Proceedings DARPA Speech and Natural Language Processing Workshop, Hidden Valley, PA, pp. 705–708. Also in Proceedings of IEEE ICASSP-91, 705-708.
Stolcke, A. (2002). Srilm - an extensible language modeling toolkit. In ICSLP-02, Denver, CO.
Tomokiyo, L. M. and Waibel, A. (2001). Adaptation methods for non-native speech. In Proceedings of Multilinguality in Spoken Language Processing, Aalborg, Denmark.
Tyler, L. K. (1984). The structure of the initial cohort: Evidence from gating. Perception & Psychophysics, 36(5), 417–427.
Wang, M. Q. and Hirschberg, J. (1992). Automatic classification of intonational phrasing boundaries. Computer Speech and Language, 6(2), 175–196.
Wang, Z., Schultz, T., and Waibel, A. (2003). Comparison of acoustic model adaptation techniques on non-native speech. In IEEE ICASSP, Vol. 1, pp. 540–543.
Ward, W. (1989). Modelling non-verbal sounds for speech recognition. In HLT '89: Proceedings of the Workshop on Speech and Natural Language, Cape Cod, MA, pp. 47–50.
Warren, R. M. (1970). Perceptual restoration of missing speech sounds. Science, 167, 392–393.
Wegmann, S., McAllaster, D., Orloff, J., and Peskin, B. (1996). Speaker normalisation on conversational telephone speech. In IEEE ICASSP-96, Atlanta, GA.
Weintraub, M., Taussig, K., Hunicke-Smith, K., and Snodgras, A. (1996). Effect of speaking style on LVCSR performance. In ICSLP-96, Philadelphia, PA, pp. 16–19.
Welling, L., Ney, H., and Kanthak, S. (2002). Speaker adaptive modeling by vocal tract normalisation. IEEE Transactions on Speech and Audio Processing, 10, 415–426.
Woodland, P. C., Leggetter, C. J., Odell, J. J., Valtchev, V., and Young, S. J. (1995). The 1994 htk large vocabulary speech recognition system. In IEEE ICASSP.
Woodland, P. C. and Povey, D. (2002). Large scale discriminative training of hidden Markov models for speech recognition. Computer Speech and Language, 16, 25–47.
Woodland, P. C. (2001). Speaker adaptation for continuous density HMMs: A review. In Juncqua, J.-C. and Wellekens, C. (Eds.), Proceedings of the ITRW 'Adaptation Methods For Speech Recognition', Sophia-Antipolis, France.
Young, S. J. (1984). Generating multiple solutions from connected word dp recognition algorithms. Proceedings of the Institute of Acoustics, 6(4), 351–354.
Young, S. J., Odell, J. J., and Woodland, P. C. (1994). Tree-based state tying for high accuracy acoustic modelling. In Proceedings ARPA Workshop on Human Language Technology, pp. 307–312.
Young, S. J., Russell, N. H., and Thornton, J. H. S. (1989). Token passing: A simple conceptual model for connected speech recognition systems. Tech. rep. CUED/F-INFENG/TR.38, Cambridge University Engineering Department, Cambridge, England.
Young, S. J. and Woodland, P. C. (1994). State clustering in HMM-based continuous speech recognition. Computer Speech and Language, 8(4), 369–394.
Young, S. J., Evermann, G., Gales, M., Hain, T., Kershaw, D., Moore, G., Odell, J. J., Ollason, D., Povey, D., Valtchev, V., and Woodland, P. C. (2005). The HTK Book. Cambridge University Engineering Department.
Zheng, Y., Sproat, R., Gu, L., Shafran, I., Zhou, H., Su, Y., Jurafsky, D., Starr, R., and Yoon, S.-Y. (2005). Accent detection
and speech recognition for shanghai-accented mandarin. In InterSpeech 2005, Lisbon, Portugal.
Zweig, G. (1998). Speech Recognition with Dynamic Bayesian Networks. Ph.D. thesis, University of California, Berkeley.
11 COMPUTATIONAL PHONOLOGY
bidakupadotigolabubidakutupiropadotigolabutupirobidaku...
Word segmentation stimulus (Saffran et al., 1996a)
Recall from Ch. 7 that phonology is the area of linguistics that describes the systematic way that sounds are differently realized in different environments, and how this system of sounds is related to the rest of the grammar. This chapter introduces computational phonology, the use of computational models in phonological theory.
One focus of computational phonology is on computational models of phonological representation, and on how to use phonological models to map from surface phonological forms to underlying phonological representation. Models in (non-computational) phonological theory are generative; the goal of the model is to represent how an underlying form can generate a surface phonological form. In computation, we are generally more interested in the alternative problem of phonological parsing; going from surface form to underlying structure. One major tool for this task is the finite-state automaton, which is employed in two families of models: finite-state phonology and optimality theory.
A related kind of phonological parsing task is syllabification: the task of assigning syllable structure to sequences of phones. Besides its theoretical interest, syllabification turns out to be a useful practical tool in aspects of speech synthesis such as pronunciation dictionary design. We therefore summarize a few practical algorithms for syllabification.
Finally, we spend the remainder of the chapter on the key problem of how phonological and morphological representations can be learned.
11.1 FINITE-STATE PHONOLOGY
Ch. 3 showed that spelling rules can be implemented by transducers. Phonological rules can be implemented as transducers in the same way; indeed the original work by Johnson (1972) and Kaplan and Kay (1981) on finite-state models was based on phonological rules rather than spelling rules. There are a number of different models of computational phonology that use finite automata in various ways to realize phonology.
logical rules. We will describe the two-level morphology of Koskenniemi (1983) first mentioned in Ch. 3. Let's begin with the intuition, by seeing the transducer in Fig. 11.1 which models the simplified flapping rule in (11.1):
$$ \mathrm{/t/}\to[\mathrm{dx}]/\mathrm{\check{V}}\mathrm{\_\-V} $$

The transducer in Fig. 11.1 accepts any string in which flaps occur in the correct places (after a stressed vowel, before an unstressed vowel), and rejects strings in which flapping doesn't occur, or in which flapping occurs in the wrong environment. $ ^{1} $
We’ve seen both transducers and rules before; the intuition of two-level morphology is to augment the rule notation to correspond more naturally to transducers. We motivate his idea by beginning with the notion of rule ordering. In a traditional phonological system, many different phonological rules apply between the lexical form and the surface form. Sometimes these rules interact; the output from one rule affects the input to another rule. One way to implement rule-interaction in a transducer system is to run transducers in a cascade. Consider, for example, the rules that are needed to deal with the phonological behavior of the English noun plural suffix -s. This suffix is pronounced [ix z] after the phones [s], [sh], [z], [zh], or [jh] (so peaches is pronounced [p iy ch ix z], and faxes is pronounced [f ae k s ix z]), [z] after voiced sounds (pigs is pronounced [p ih g z]), and [s] after unvoiced sounds (cats is pronounced [k ae t s]). We model this variation by writing phonological rules for the realization of the morpheme in different contexts. We first need to choose one of these three forms ([s], [z], [ix z]) as the “lexical” pronunciation of the suffix; we chose [z] only because it turns out to simplify rule writing. Next we write two phonological rules. One, similar to the E-insertion spelling rule of page ??, inserts an [ix] after a morpheme-final sibilant and before the plural morpheme [z]. The other makes sure that the -s suffix is properly realized as [s] after unvoiced consonants.
$$ \epsilon~\rightarrow~ix~/~[+\mathrm{s i b i l a n t}]~\overset{\frown}{~_\_\ }\mathrm{~z~}\# $$
(11.3)
$$ \mathrm{~z~}\to\mathrm{~s~}/\left[\mathrm{-voice}\right]^{\wedge}\_\_\# $$
These two rules must be ordered; rule (11.2) must apply before (11.3). This is because the environment of (11.2) includes z, and the rule (11.3) changes z. Consider running both rules on the lexical form fox concatenated with the plural -s:
$$ \textit{Lexical form:}\qquad\texttt{f a a k}\emph{z} $$
$$ (11.2)\;a p p l i e s{:}\qquad\textrm{f a a k s}\hat{\textrm{a i x}}z $$
$$ (11.3)does n^{\prime}t~a p p l y\text{:f}a a k s^{\prime}i x z $$
If the devoicing rule (11.3) was ordered first, we would get the wrong result. This situation, in which one rule destroys the environment for another, is called bleeding: $ ^{2} $
$$ \textit{Lexical form:}\qquad\texttt{f a a k s}^{\wedge}z $$
$$ (11.3)a p p l i e s:\quad\mathrm{~f~a~a~k~s~}\mathrm{~s~} $$
$$ (11.2)does n^{\prime}t~a p p l y\text{:f}a a k s\hat{}\text{s} $$
As was suggested in Ch. 3, each of these rules can be represented by a transducer. Since the rules are ordered, the transducers would also need to be ordered. For example if they are placed in a $ \text{cascade} $, the output of the first transducer would feed the input of the second transducer.
Many rules can be cascaded together this way. As Ch. 3 discussed, running a cascade, particularly one with many levels, can be unwieldy, and so transducer cascades are usually replaced with a single more complex transducer by composing the individual transducers.
Koskenniemi's method of two-level morphology that was sketchily introduced in Ch. 3 is another way to solve the problem of rule ordering. Koskenniemi (1983) observed that most phonological rules in a grammar are independent of one another; that feeding and bleeding relations between rules are not the norm. $ ^{3} $ Since this is the case, Koskenniemi proposed that phonological rules be run in parallel rather than in series. The cases where there is rule interaction (feeding or bleeding) we deal with by slightly modifying some rules. Koskenniemi's two-level rules can be thought of as a way of expressing declarative constraints on the well-formedness of the lexical-surface mapping.
Two-level rules also differ from traditional phonological rules by explicitly coding when they are obligatory or optional, by using four differing rule operators; the ⇔ rule corresponds to traditional obligatory phonological rules, while the ⇒ rule implements optional rules:
| Rule type | Interpretation |
| --- | --- |
| a:b \Leftarrow c $ \underline{\text{d}} $ | a is always realized as b in the context c $ \underline{\text{d}} $ |
| a:b \Rightarrow c $ \underline{\text{d}} $ | a may be realized as b only in the context c $ \underline{\text{d}} $ |
| a:b \Leftrightarrow c $ \underline{\text{d}} $ | a must be realized as b in context c $ \underline{\text{d}} $ and nowhere else |
| a:b / \Leftarrow c $ \underline{\text{d}} $ | a is never realized as b in the context c $ \underline{\text{d}} $ |
The most important intuition of the two-level rules, and the mechanism that lets them avoid feeding and bleeding, is their ability to represent constraints on two levels. This is based on the use of the colon (“:”), which was touched on very briefly in Ch. 3. The symbol $ a:b $ means a lexical $ a $ that maps to a surface $ b $. Thus $ a:b \Leftrightarrow :c $ ___ means $ a $ is realized as $ b $ after a surface $ c $. By contrast $ a:b \Leftrightarrow c $: ___ means that $ a $ is realized as $ b $ after a lexical $ c $. As discussed in Ch. 3, the symbol $ c $ with no colon is equivalent to $ c:c $ that means a lexical $ c $ which maps to a surface $ c $.
Fig. 11.2 shows an intuition for how the two-level approach avoids ordering for the ix-insertion and z-devoicing rules. The idea is that the z-devoicing rule maps a lexical z-insertion to a surface s and the ix rule refers to the lexical z.

The two-level rules that model this constraint are shown in (11.4) and (11.5):
$$ \epsilon:\mathrm{i x}\iff\left[\mathrm{+s i b i l a n t}\right]:\hat{\mathbf{\textit{a m z}}}\;\underline{\quad}\;\mathrm{z}:\;\# $$
$$ \mathrm{z:\text{s} \quad \Leftrightarrow~[-voice]:~^~ \quad \underline{\quad}~\#} $$
As Ch. 3 discussed, there are compilation algorithms for creating automata from rules. Kaplan and Kay (1994) give the general derivation of these algorithms, and Antworth (1990) gives one that is specific to two-level rules. The automata corresponding to the two rules are shown in Fig. 11.3 and Fig. 11.4. Fig. 11.3 is based on Figure 3.14 of Ch. 3; see page 78 for a reminder of how this automaton works. Note in Fig. 11.3 that the plural morpheme is represented by z, indicating that the constraint is expressed about a lexical rather than surface z.
Fig. 11.5 shows the two automata run in parallel on the input [f aa k s^z]. Note that both the automata assumes the default mapping ^:ε to remove the morpheme boundary, and that both automata end in an accepting state.