← 学习库 Speech and Language Processing 本册目录

8.4.2 Diphone concatenation and TD-PSOLA for prosodic adjustment

We are now ready to see the remaining steps for synthesizing an individual utterance. Assume that we have completed text analysis for the utterance, and hence arrived at a sequence of diphones and prosodic targets, and that we have also grabbed the appropriate sequence of diphones from the diphone database. Next we need to concatenate the diphones together and then adjust the prosody (pitch, energy, and duration) of the diphone sequence to match the prosodic requirements from the intermediate representation.

Given two diphones, what do we need to do to concatenate them successfully? If the waveforms of the two diphones edges across the juncture are very different, a perceptible click will result. Thus we need to apply a windowing function to the edge of both diphones so that the samples at the juncture have low or zero amplitude. Furthermore, if both diphones are voiced, we need to ensure that the two diphones are joined pitch-synchronously. This means that the pitch periods at the end of the first diphone must line up with the pitch periods at the beginning of the second diphone; otherwise the resulting single irregular pitch period at the juncture is perceptible as well.

Now given our sequence of concatenated diphones, how do we modify the pitch and duration to meet our prosodic requirements? It turns out there is a very simple algorithm for doing this called TD-PSOLA (Time-Domain Pitch-Synchronous OverLap-and-Add).

As we just said, a pitch-synchronous algorithm is one in which we do something at each pitch period or epoch. For such algorithms it is important to have very accurate pitch markings: measurements of exactly where each pitch pulse or epoch occurs. An epoch can be defined by the instant of maximum glottal pressure, or alternatively by the instant of glottal closure. Note the distinction between pitch marking or epoch detection and pitch tracking. Pitch tracking gives the value of F0 (the average cycles per second of the glottis) at each particular point in time, averaged over a neighborhood.

原书第 302 页

Pitch marking finds the exact point in time at each vibratory cycle at which the vocal folds reach some specific point (epoch).

Epoch-labeling can be done in two ways. The traditional way, and still the most accurate, is to use an electroglottograph or EGG (electroglottograph) (often also called a laryngograph or Lx (laryngograph)). An EGG is a device which straps onto the (outside of the) speaker's neck near the larynx and sends a small current through the Adam's apple. A transducer detects whether the glottis is open or closed by measuring the impedance across the vocal folds. Some modern synthesis databases are still recorded with an EGG. The problem with using an EGG is that it must be attached to the speaker while they are recording the database. Although an EGG isn't particularly invasive, this is still annoying, and the EGG must be used during recording; it can't be used to pitch-mark speech that has already been collected. Modern epoch detectors are now approaching a level of accuracy that EGGs are no longer used in most commercial TTS engines. Algorithms for epoch detection include Brookes and Loke (1999), Veldhuis (2000).

Given an epoch-labeled corpus, the intuition of TD-PSOLA is that we can modify the pitch and duration of a waveform by extracting a frame for each pitch period (windowed so that the frame doesn't have sharp edges) and then recombining these frames in various ways by simply overlapping and adding the windowed pitch period frames (we will introduce the idea of windows in Sec. ??). The idea that we modify a signal by extracting frames, manipulating them in some way and then recombining them by adding up the overlapped signals is called the overlap-and-add or OLA algorithm; TD-PSOLA is a special case of overlap-and-add in which the frames are pitch-synchronous, and the whole process takes place in the time domain.

For example, in order to assign a specific duration to a diphone, we might want to lengthen the recorded master diphone. To lengthen a signal with TD-PSOLA, we simply insert extra copies of some of the pitch-synchronous frames, essentially duplicating a piece of the signal. Fig. 8.16 shows the intuition.

TD-PSOLA can also be used to change the F0 value of a recorded diphone to give a higher or lower value. To increase the F0, we extract each pitch-synchronous frame from the original recorded diphone signal, place the frames closer together (overlapping them), with the amount of overlap determined by the desired period and hence frequency, and then add up the overlapping signals to produce the final signal. But note that by moving all the frames closer together, we make the signal shorter in time! Thus in order to change the pitch while holding the duration constant, we need to add duplicate frames.

Fig. 8.17 shows the intuition; in this figure we have explicitly shown the extracted pitch-synchronous frames which are overlapped and added; note that the frames moved closer together (increasing the pitch) while extra frames have been added to hold the duration constant.

原书第 303 页
Image
Figure 8.16 TD-PSOLA for duration modification. Individual pitch-synchronous frames can be duplicated to lengthen the signal (as shown here), or deleted to shorten the signal.

8.5 UNIT SELECTION (WAVEFORM) SYNTHESIS

Diphone waveform synthesis suffers from two main problems. First, the stored di-phone database must be modified by signal process methods like PSOLA to produce the desired prosody. Any kind of signal processing of the stored speech leaves artifacts in the speech which can make the speech sound unnatural. Second, diphone synthesis only captures the coarticulation due to a single neighboring phone. But there are many more global effects on phonetic realization, including more distant phones, syllable structure, the stress patterns of nearby phones, and even word-level effects.

For this reason, modern commercial synthesizers are based on a generalization of diphone synthesis called unit selection synthesis. Like diphone synthesis, unit selection synthesis is a kind of concatenative synthesis algorithm. It differs from classic diphone synthesis in two ways:

1. In diphone synthesis the database stores exactly one copy of each diphone, while in unit selection, the unit database is many hours long, containing many copies of each diphone.

2. In diphone synthesis, the prosody of the concatenated units is modified by PSOLA or similar algorithms, while in unit selection no (or minimal) signal processing is applied to the concatenated units.

The strengths of unit selection are due to the large unit database. In a sufficiently large database, entire words or phrases of the utterance we want to synthesize may be already present in the database, resulting in an extremely natural waveform for these

原书第 304 页
Image
Figure 8.17 TD-PSOLA for pitch (F0) modification. In order to increase the pitch, the individual pitch-synchronous frames are extracted, Hanning windowed, moved closer together and then added up. To decrease the pitch, we move the frames further apart. Increasing the pitch will result in a shorter signal (since the frames are closer together), so we also need to duplicate frames if we want to change the pitch while holding the duration constant.

words or phrases. In addition, in cases where we can't find a large chunk and have to back off to individual diphones, the fact that there are so many copies of each diphone makes it more likely that we will find one that will fit in very naturally.

The architecture of unit selection can be summarized as follows. We are given a large database of units; let's assume these are diphones (although it's also possible to do unit selection with other kinds of units such half-phones, syllables, or half-syllables). We are also given a characterization of the target 'internal representation', i.e. a phone string together with features such as stress values, word identity, F0 information, as described in Fig. 8.1.

The goal of the synthesizer is to select from the database the best sequence of

原书第 305 页

diphone units that corresponds to the target representation. What do we mean by the 'best' sequence? Intuitively, the best sequence would be one in which:

each diphone unit we select exactly meets the specifications of the target diphone (in terms of F0, stress level, phonetic neighbors, etc)

each diphone unit concatenates smoothly with its neighboring units, with no perceptible break.

Of course, in practice, we can't guarantee that there will be a unit which exactly meets our specifications, and we are unlikely to find a sequence of units in which every single join is imperceptible. Thus in practice unit selection algorithms implement a gradient version of these constraints, and attempt to find the sequence of unit which at least minimizes these two costs:

target cost $ T(u_t, s_t) $: how well the target specification $ s_t $ matches the potential unit $ u_t $ join cost $ J(u_t, u_{t+1}) $: how well (perceptually) the potential unit $ u_t $ joins with its potential neighbor $ u_{t+1} $

The T and J values are expressed as costs meaning that high values indicate bad matches and bad joins (Hunt and Black, 1996a).

Formally, then, the task of unit selection synthesis, given a sequence S of T target specifications, is to find the sequence $ \hat{U} $ of T units from the database which minimizes the sum of these costs:

$$ \hat{U}=\underset{U}{\operatorname{a r g m i n}}\sum_{t=1}^{T}T(s_{t},u_{t})+\sum_{t=1}^{T-1}J(u_{t},u_{t+1}) $$

Let's now define the target cost and the join cost in more detail before we turn to the decoding and training tasks.

The target cost measures how well the unit matches the target diphone specification. We can think of the specification for each diphone target as a feature vector; here are three sample vectors for three target diphone specifications, using dimensions (features) like should the syllable be stressed, and where in the intonational phrase should the diphone come from:

$$ \begin{array}{l l l l l}{{{/i h\_{t}/,+s t r e s s,p h r a s e\_{i} n t e r n a l,h i g h\_{F} O,c o n t e n t\_{w} o r d}}}&{{{w o r d}}}\\ {{/n\_{t}/,-s t r e s s,p h r a s e\_{f} i n a l,h i g h\_{F} O,f u n c t i o n\_{w} o r d}}&{{{w o r d}}}\\ {{/d h\_{a} x/,-s t r e s s,p h r a s e\_{i} n i t i a l,l o w\_{F} O,w o r d\setminus\mathrm{t h e^{\prime}}}}\end{array} $$

We’d like the distance between the target specification $s$ and the unit to be some function of the $h$ now different from the unit is on each of these dimensions from the specification. Let’s assume that for each dimension $p$, we can come up with some subcost $T_p(s_t[p], u_j[p])$. The subcost for a binary feature like $p$ is $1$ or 0. The subcost for a continuous feature like $F_0$ might be the difference (or $\log$ difference) between the specification $F_0$ and unit $F_0$. Since some dimensions are more important to speech perceptions than others, we’ll also want to weight each dimension. The simplest way to combine all these subcosts is just to assume that they are independent and additive. Using this model, the total target cost for a given target/unit pair is the weighted sum over all these subcosts for each feature/dimension:

原书第 306 页

$$ T(s_{t},u_{j})=\sum_{p=1}^{P}w_{p}T_{p}(s_{t}[p],u_{j}[p]) $$

The target cost is a function of the desired diphone specification and a unit from the database. The join cost, by contrast, is a function of two units from the database. The goal of the join cost is to be low (0) when the join is completely natural, and high when the join would be perceptible or jarring. We do this by measuring the acoustic similarity of the edges of the two units that we will be joining. If the two units have very similar energy, F0, and spectral features, they will probably join well. Thus as with the target cost, we compute a join cost by summing weighted subcosts:

$$ J(u_{t},u_{t+1})=\sum_{p=1}^{P}w_{p} J_{p}(u_{t}[p],u_{t+1}[p]) $$

The three subcosts used in the classic Hunt and Black (1996b) algorithm are the cepstral distance at the point of concatenation, and the absolute differences in log power and F0. We will introduce the cepstrum in Sec. ??.

In addition, if the two units $ u_t $ and $ u_{t+1} $ to be concatenated were consecutive diphones in the unit database (i.e. they followed each other in the original utterance), then we set the join cost to 0: $ J(u_t, u_{t+1}) = 0 $. This is an important feature of unit selection synthesis, since it encourages large natural sequences of units to be selected from the database.

How do we find the best sequence of units which minimizes the sum of the target and join costs as expressed in Eq. 8.21? The standard method is to think of the unit selection problem as a Hidden Markov Model. The target units are the observed outputs, and the units in the database are the hidden states. Our job is to find the best hidden state sequence. We will use the Viterbi algorithm to solve this problem, just as we saw it in Ch. 5 and Ch. 6, and will see it again in Ch. 9. Fig. 8.18 shows a sketch of the search space as well as the best (Viterbi) path that determines the best unit sequence.

The weights for join and target costs are often set by hand, since the number of weights is small (on the order of 20) and machine learning algorithms don't always achieve human performance. The system designer listens to entire sentences produced by the system, and chooses values for weights that result in reasonable sounding utterances. Various automatic weight-setting algorithms do exist, however. Many of these assume we have some sort of distance function between the acoustics of two sentences, perhaps based on cepstral distance. The method of Hunt and Black (1996b), for example, holds out a test set of sentences from the unit selection database. For each of these test sentences, we take the word sequence and synthesize a sentence waveform (using units from the other sentences in the training database). Now we compare the acoustics of the synthesized sentence with the acoustics of the true human sentence. Now we have a sequence of synthesized sentences, each one associated with a distance function to its human counterpart. Now we use linear regression based on these distances to set the target cost weights so as to minimize the distance.

There are also more advanced methods of assigning both target and join costs. For example, above we computed target costs between two units by looking at the features

原书第 307 页
Image
Figure 8.18 The process of decoding in unit selection. The figure shows the sequence of target (specification) diphones for the word six, and the set of possible database diphone units that we must search through. The best (Viterbi) path that minimizes the sum of the target and join costs is shown in bold.

of the two units, doing a weighted sum of feature costs, and choosing the lowest-cost unit. An alternative approach (which the new reader might need to come back to after learning the speech recognition techniques introduced in the next chapters) is to map the target unit into some acoustic space, and then find a unit which is near the target in that acoustic space. In the method of Donovan and Eide (1998), Donovan and Woodland (1995), for example, all the training units are clustered using the decision tree algorithm of speech recognition described in Sec. ??. The decision tree is based on the same features described above, but here for each set of features, we follow a path down the decision tree to a leaf node which contains a cluster of units that have those features. This cluster of units can be parameterized by a Gaussian model, just as for speech recognition, so that we can map a set of features into a probability distribution over cepstral values, and hence easily compute a distance between the target and a unit in the database. As for join costs, more sophisticated metrics make use of how perceivable a particular join might be (Wouters and Macon, 1998; Syrdal and Konkie, 2004; Bulyko and Ostendorf, 2001).

8.6 EVALUATION

Speech synthesis systems are evaluated by human listeners. The development of a good automatic metric for synthesis evaluation, that would eliminate the need for expensive and time-consuming human listening experiments, remains an open and existing research topic.

The minimal evaluation metric for speech synthesis systems is intelligibility: the ability of a human listener to correctly interpret the words and meaning of the synthesized utterance. A further metric is quality; an abstract measure of the naturalness, fluency, or clarity of the speech.

原书第 308 页

The most local measures of intelligibility test the ability of a listener to discriminate between two phones. The Diagnostic Rhyme Test (DRT) (Voiers et al., 1975) tests the intelligibility of initial consonants. It is based on 96 pairs of confusable rhyming words which differ only in a single phonetic feature, such as (dense/tense) or bond/pond (differing in voicing) or mean/beat or neck/deck (differing in nasality), and so on. For each pair, listeners hear one member of the pair, and indicate which they think it is. The percentage of right answers is then used as an intelligibility metric. The Modified Rhyme Test (MRT) (House et al., 1965) is a similar test based on a different set of 300 words, consisting of 50 sets of 6 words. Each 6-word set differs in either initial or final consonants (e.g., went, sent, bent, dent, tent, rent or bat, bad, back, bass, ban, bath). Listeners are again given a single word and must identify from a closed list of six words; the percentage of correct identifications is again used as an intelligibility metric.

Since context effects are very important, both DRT and MRT words are embedded in carrier phrases like the following:

Now we will say again.

In order to test larger units than single phones, we can use semantically unpredictable sentences (SUS) (Benoît et al., 1996). These are sentences constructed by taking a simple POS template like DET ADJ NOUN VERB DET NOUN and inserting random English words in the slots, to produce sentences like

The unsure steaks closed the fish.

Measures of intelligibility like DRT/MRT and SUS are designed to factor out the role of context in measuring intelligibility. While this allows us to get a carefully controlled measure of a system's intelligibility, such as contextual or semantically unpredictable sentences aren't a good fit to how TTS is used in most commercial applications. Thus in commercial applications instead of DRT or SUS, we generally test intelligibility using situations that mimic the desired applications; reading addresses out loud, reading lines of news text, and so on.

To further evaluate the quality of the synthesized utterances, we can play a sentence for a listener and ask them to give a mean opinion score (MOS), a rating of how good the synthesized utterances are, usually on a scale from 1-5. We can then compare systems by comparing their MOS scores on the same sentences (using, e.g., t-tests to test for significant differences).

If we are comparing exactly two systems (perhaps to see if a particular change actually improved the system), we can use AB tests in AB tests, we play the same sentence synthesized by two different systems (an A and a B system). The human listener chooses which of the two utterances they like better. We can do this for 50 sentences and compare the number of sentences preferred for each system. In order to avoid ordering preferences, for each sentence we must present the two synthesized waveforms in random order.

原书第 309 页

BIBLIOGRAPHICAL AND HISTORICAL NOTES

As we noted at the beginning of the chapter, speech synthesis is one of the earliest fields of speech and language processing. The 18th century saw a number of physical models of the articulation process, including the von Kempenen model mentioned above, as well as the 1773 vowel model of Kratzenstein in Copenhagen using organ pipes.

But the modern era of speech synthesis can clearly be said to have arrived by the early 1950's, when all three of the major paradigms of waveform synthesis had been proposed (formant synthesis, articulatory synthesis, and concatenative synthesis).

Concatenative synthesis seems to have been first proposed by Harris (1953) at Bell Laboratories, who literally spliced together pieces of magnetic tape corresponding to phones. Harris's proposal was actually more like unit selection synthesis than diphone synthesis, in that he proposed storing multiple copies of each phone, and proposed the use of a join cost (choosing the unit with the smoothest formant transitions with the neighboring unit). Harris's model was based on the phone, rather than diphone, resulting in problems due to coarticulation. Peterson et al. (1958) added many of the basic ideas of unit selection synthesis, including the use of diphones, a database with multiple copies of each diphone with differing prosody, and each unit labeled with intonational features including F0, stress, and duration, and the use of join costs based on F0 and formant distant between neighboring units. They also proposed microconcatenation techniques like windowing the waveforms. The Peterson et al. (1958) model was purely theoretical, however, and concatenative synthesis was not implemented until the 1960's and 1970's, when diphone synthesis was first implemented (Dixon and Maxey, 1968; Olive, 1977). Later diphone systems included larger units such as consonant clusters (Olive and Liberman, 1979). Modern unit selection, including the idea of large units of non-uniform length, and the use of a target cost, was invented by Sagisaka (1988), Sagisaka et al. (1992). Hunt and Black (1996b) formalized the model, and put it in the form in which we have presented it in this chapter in the context of the ATR CHATR system (Black and Taylor, 1994). The idea of automatically generating synthesis units by clustering was first invented by Nakajima and Hamada (1988), but was developed mainly by (Donovan, 1996) by incorporating decision tree clustering algorithms from speech recognition. Many unit selection innovations took place as part of the ATT NextGen synthesizer (Syrdal et al., 2000; Syrdal and Conkie, 2004).

We have focused in this chapter on concatenative synthesis, but there are two other paradigms for synthesis: \textit{formant synthesis}, in which we attempt to build rules which generate artificial spectra, including especially formants, and \textit{articulatory synthesis}, in which we attempt to directly model the physics of the vocal tract and articulatory process.

Formant synthesizers originally were inspired by attempts to mimic human speech by generating artificial spectrograms. The Haskins Laboratories Pattern Playback Machine generated a sound wave by painting spectrogram patterns on a moving transparent belt, and using reflectance to filter the harmonics of a waveform (Cooper et al., 1951); other very early formant synthesizers include Lawrence (1953) and Fant (3951). Perhaps the most well-known of the formant synthesizers were the Klatt formant syn-

原书第 310 页

thesizer and its successor systems, including the MITalk system (Allen et al., 1987), and the Klattalk software used in Digital Equipment Corporation's DECtalk (Klatt, 1982). See Klatt (1975) for details.

Articulatory synthesizers attempt to synthesize speech by modeling the physics of the vocal tract as an open tube. Representative models, both early and somewhat more recent include Stevens et al. (1953), Flanagan et al. (1975), Fant (1986) See Klatt (1975) and Flanagan (1972) for more details.

Development of the text analysis components of TTS came somewhat later, as techniques were borrowed from other areas of natural language processing. The input to early synthesis systems was not text, but rather phonemes (typed in on punched cards). The first text-to-speech system to take text as input seems to have been the system of Umeda and Teranishi (Umeda et al., 1968; Teranishi and Umeda, 1968; Umeda, 1976). The system included a lexicalized parser which was used to assign prosodic boundaries, as well as accent and stress; the extensions in Coker et al. (1973) added additional rules, for example for deaccenting light verbs and explored articulatory models as well. These early TTS systems used a pronunciation dictionary for word pronunciations. In order to expand to larger vocabularies, early formant-based TTS systems such as MITlak (Allen et al., 1987) used letter-to-sound rules instead of a dictionary, since computer memory was far too expensive to store large dictionaries.

Modern grapheme-to-phoneme models derive from the influential early probabilistic grapheme-to-phoneme model of Lucassen and Mercer (1984), which was originally proposed in the context of speech recognition. The widespread use of such machine learning models was delayed, however, because early anecdotal evidence suggested that hand-written rules worked better than e.g., the neural networks of Sejnowski and Rosenberg (1987). The careful comparisons of Damper et al. (1999) showed that machine learning methods were in generally superior. A number of such models make use of pronunciation by analogy (Byrd and Chodorow, 1985; ?; Daelemans and van den Bosch, 1997; Marchand and Damper, 2000) or latent analogy (Bellegarda, 2005); HMMs (Taylor, 2005) have also been proposed. The most recent work makes use of joint graphene models, in which the hidden variables are phoneme-grapheme pairs and the probabilistic model is based on joint rather than conditional likelihood (Deligne et al., 1995; Luk and Damper, 1996; Galescu and Allen, 2001; Bisani and Ney, 2002; Chen, 2003).

There is a vast literature on prosody. Besides the ToBI and TILT models described above, other important computational models include the Fujisaki model (Fujisaki and Ohno, 1997). IVIe (Grabe, 2001) is an extension of ToBI that focuses on labelling different varieties of English (Grabe et al., 2000). There is also much debate on the units of intonational structure (intonational phrases (Beckman and Pierrehumbert, 1986), intonation units (Du Bois et al., 1983) or tone units (Crystal, 1969)), and their relation to clauses and other syntactic units (Chomsky and Halle, 1968; Langendoen, 1975; Streeter, 1978; Hirschberg and Pierrehumbert, 1986; Selkirk, 1986; Nespor and Vogel, 1986; Croft, 1995; Ladd, 1996; Ford and Thompson, 1996; Ford et al., 1996).

One of the most exciting new paradigms for speech synthesis is HMM synthesis, first proposed by Tokuda et al. (1995b) and elaborated in Tokuda et al. (1995a), Tokuda et al. (2000), and Tokuda et al. (2003). See also the textbook summary of HMM synthesis in Taylor (2008).

原书第 311 页

BLIZZARD CHALLENGE

More details on TTS evaluation can be found in Huang et al. (2001) and Gibbon et al. (2000). Other descriptions of evaluation can be found in the annual speech synthesis competition called the Blizzard Challenge (Black and Tokuda, 2005; Bennett, 2005).

Much recent work on speech synthesis has focused on generating emotional speech (Cahn, 1990; Bulut1 et al., 2002; Hamza et al., 2004; Eide et al., 2004; Lee et al., 2006; Schroder, 2006, inter alia)

Two classic text-to-speech synthesis systems are described in Allen et al. (1987) (the MITalk system) and Sproat (1998b) (the Bell Labs system). Recent textbooks include Dutoit (1997), Huang et al. (2001), Taylor (2008), and Alan Black's online lecture notes at http://festvox.org/festtut/notes/festtut_toc.html. Influential collections of papers include van Santen et al. (1997), Sagisaka et al. (1997), Narayanan and Alwan (2004). Conference publications appear in the main speech engineering conferences (INTERSPEECH, IEEE ICASSP), and the Speech Synthesis Workshops. Journals include Speech Communication, Computer Speech and Language, the IEEE Transactions on Audio, Speech, and Language Processing, and the ACM Transactions on Speech and Language Processing.

EXERCISES

8.1 Implement the text normalization routine that deals with MONEY, i.e. mapping strings of dollar amounts like $45, $320, and $4100 to words (either writing code directly or designing an FST). If there are multiple ways to pronounce a number you may pick your favorite way.

8.2 Implement the text normalization routine that deals with NTEL, i.e. seven-digit phone numbers like 555-1212, 555-1300, and so on. You should use a combination of the paired and trailing unit methods of pronunciation for the last four digits. (Again you may either write code or design an FST).

8.3 Implement the text normalization routine that deals with type DATE in Fig. 8.4

8.4 Implement the text normalization routine that deals with type NTIME in Fig. 8.4.

8.5 (Suggested by Alan Black). Download the free Festival speech synthesizer. Augment the lexicon to correctly pronounce the names of everyone in your class.

8.6 Download the Festival synthesizer. Record and train a diphone synthesizer using your own voice.

原书第 312 页

Allen, J., Hunnicut, M. S., and Klatt, D. H. (1987). From Text to Speech: The MITalk system. Cambridge University Press.

Anderson, M. J., Pierrehumbert, J. B., and Liberman, M. Y. (1984). Improving intonational phrasing with syntactic information. In IEEE ICASSP-84, pp. 2.8.1–2.8.4.

Bachenko, J. and Fitzpatrick, E. (1990). A computational grammar of discourse-neutral prosodic phrasing in English. Computational Linguistics, 16(3), 155–170.

Beckman, M. E. and Ayers, G. M. (1997). Guidelines for ToBI labelling..

Beckman, M. E. and Hirschberg, J. (1994). The tobi annotation conventions. Manuscript, Ohio State University.

Beckman, M. E. and Pierrehumbert, J. B. (1986). Intonational structure in English and Japanese. Phonology Yearbook, 3, 255–310.

Bellegarda, J. R. (2005). Unsupervised, language-independent grapheme-to-phoneme conversion by latent analogy. Speech Communication, 46(2), 140–152.

Bennett, C. (2005). Large scale evaluation of corpus-based synthesizers: Results and lessons from the blizzard challenge 2005. In EUROSPEECH-05.

Benoît, C., Grice, M., and Hazan, V. (1996). The SUS test: A method for the assessment of text-to-speech synthesis intelligibility using Semantically Unpredictable Sentences. Speech Communication, 18(4), 381–392.

Bisani, M. and Ney, H. (2002). Investigations on joint-multigram models for grapheme-to-phoneme conversion. In ICSLP-02, Vol. 1, pp. 105–108.

Black, A. W. and Taylor, P. (1994). CHATR: a generic speech synthesis system. In COLING-94, Kyoto, Vol. II, pp. 983–986.

Black, A. W. and Hunt, A. J. (1996). Generating F0 contours from ToBI labels using linear regression. In ICSLP-96, Vol. 3, pp. 1385–1388.

Black, A. W., Lenzo, K., and Pagel, V. (1998). Issues in building general letter to sound rules. In 3rd ESCA Workshop on Speech Synthesis, Jenolan Caves, Australia.

Black, A. W., Taylor, P., and Caley, R. (1996-1999). The Festival Speech Synthesis System system. Manual and source code available at www.cstr.ed.ac.uk/projects/festival.html.

Black, A. W. and Tokuda, K. (2005). The Blizzard Challenge—2005: Evaluating corpus-based speech synthesis on common datasets. In EUROSPEECH-05.

Bolinger, D. (1972). Accent is predictable (if you're a mind-reader). Language, 48(3), 633–644.

Brookes, D. M. and Loke, H. P. (1999). Modelling energy flow in the vocal tract with applications to glottal closure and opening detection. In IEEE ICASSP-99, pp. 213–216.

Bulut1, M., Narayanan, S. S., and Syrdal, A. K. (2002). Expressive speech synthesis using a concatenative synthesizer. In ICSLP-02.

Bulyko, I. and Ostendorf, M. (2001). Unit selection for speech synthesis using splicing costs with weighted finite state transducers. In EUROSPEECH-01, Vol. 2, pp. 987–990.

Byrd, R. J. and Chodorow, M. S. (1985). Using an On-Line dictionary to find rhyming words and pronunciations for unknown words. In ACL-85, pp. 277–283.

Cahn, J. E. (1990). The generation of affect in synthesized speech. In Journal of the American Voice I/O Society, Vol. 8, pp. 1–19.

Chen, S. F. (2003). Conditional and joint models for grapheme-to-phoneme conversion. In EUROSPEECH-03.

Chomsky, N. and Halle, M. (1968). The Sound Pattern of English. Harper and Row.

CMU (1993). The Carnegie Mellon Pronouncing Dictionary v0.1. Carnegie Mellon University.

Coker, C., Umeda, N., and Browman, C. (1973). Automatic synthesis from ordinary english test. IEEE Transactions on Audio and Electroacoustics, 21(3), 293–298.

Collins, M. (1997). Three generative, lexicalised models for statistical parsing. In ACL/EACL-97, Madrid, Spain, pp. 16–23.

Conkie, A. and Isard, S. (1996). Optimal coupling of diphones. In van Santen, J. P. H., Sproat, R., Olive, J. P., and Hirschberg, J. (Eds.), Progress in Speech Synthesis. Springer.

Cooper, F. S., Liberman, A. M., and Borst, J. M. (1951). The Interconversion of Audible and Visible Patterns as a Basis for Research in the Perception of Speech. Proceedings of the National Academy of Sciences, 37(5), 318–325.

Croft, W. (1995). Intonation units and grammatical structure. Linguistics, 33, 839–882.

Crystal, D. (1969). Prosodic systems and intonation in English. Cambridge University Press.

Daelemans, W. and van den Bosch, A. (1997). Language-independent data-oriented grapheme-to-phoneme conversion. In van Santen, J. P. H., Sproat, R., Olive, J. P., and Hirschberg, J. (Eds.), Progress in Speech Synthesis, pp. 77–89. Springer.

Damper, R. I., Marchand, Y., Adamson, M. J., and Gustafson, K. (1999). Evaluating the pronunciation component of text-to-speech systems for english: A performance comparison of different approaches. Computer Speech and Language, 13(2), 155–176.

Deligne, S., Yvon, F., and Bimbot, F. (1995). Variable-length sequence matching for phonetic transcription using joint multigrams. In EUROSPEECH-95, Madrid.

Demberg, V. (2006). Letter-to-phoneme conversion for a german text-to-speech system. Diplomarbeit Nr. 47, Universität Stuttgart.

Divay, M. and Vitale, A. J. (1997). Algorithms for grapheme-phoneme translation for English and French: Applications for database searches and speech synthesis. Computational Linguistics, 23(4), 495–523.

原书第 313 页

Dixon, N. and Maxey, H. (1968). Terminal analog synthesis of continuous speech using the diphone method of segment assembly. IEEE Transactions on Audio and Electroacoustics, 16(1), 40–50.

Donovan, R. E. (1996). Trainable Speech Synthesis. Ph.D. thesis, Cambridge University Engineering Department.

Donovan, R. E. and Eide, E. M. (1998). The IBM trainable speech synthesis system. In ICSLP-98, Sydney.

Donovan, R. E. and Woodland, P. C. (1995). Improvements in an HMM-based speech synthesiser. In EUROSPEECH-95, Madrid, Vol. 1, pp. 573–576.

Du Bois, J. W., Schuetze-Coburn, S., Cumming, S., and Paolino, D. (1983). Outline of discourse transcription. In Edwards, J. A. and Lampert, M. D. (Eds.), Talking Data: Transcription and Coding in Discourse Research, pp. 45–89. Lawrence Erlbaum.

Dutoit, T. (1997). An Introduction to Text to Speech Synthesis. Kluwer.

Eide, E. M., Bakis, R., Hamza, W., and Pitrelli, J. F. (2004). Towards synthesizing expressive speech. In Narayanan, S. S. and Alwan, A. (Eds.), Text to Speech Synthesis: New paradigms and Advances. Prentice Hall.

Fackrell, J. and Skut, W. (2004). Improving pronunciation dictionary coverage of names by modelling spelling variation. In Proceedings of the 5th Speech Synthesis Workshop.

Fant, C. G. M. (3951). Speech communication research. Ing. Vetenskaps Akad. Stockholm, Sweden, 24, 331–337.

Fant, G. M. (1986). Glottal flow: Models and interaction. Journal of Phonetics, 14, 393–399.

Fitt, S. (2002). Unisyn lexicon. http://www.cstr.ed.ac.uk/projects/unisyn/.

Flanagan, J. L., Ishizaka, K., and Shipley, K. L. (1975). Synthesis of speech from a dynamic model of the vocal cords and vocal tract. The Bell System Technical Journal, 54(3), 485–506.

Flanagan, J. L. (1972). Speech Analysis, Synthesis, and Perception. Springer.

Ford, C., Fox, B., and Thompson, S. A. (1996). Practices in the construction of turns. Pragmatics, 6, 427–454.

Ford, C. and Thompson, S. A. (1996). Interactional units in conversation: syntactic, intonational, and pragmatic resources for the management of turns. In Ochs, E., Schegloff, E. A., and Thompson, S. A. (Eds.), Interaction and Grammar, pp. 134–184. Cambridge University Press.

Fujisaki, H. and Ohno, S. (1997). Comparison and assessment of models in the study of fundamental frequency contours of speech. In ESCA workshop on Intonation: Theory Models and Applications.

Galescu, L. and Allen, J. (2001). Bi-directional conversion between graphemes and phonemes using a joint N-gram model. In Proceedings of the 4th ISCA Tutorial and Research Workshop on Speech Synthesis.

Gee, J. P. and Grosjean, F. (1983). Performance structures: A psycholinguistic and linguistic appraisal. Cognitive Psychology, 15, 411–458.

Gibbon, D., Mertins, I., and Moore, R. (2000). Handbook of Multimodal and Spoken Dialogue Systems: Resources, Terminology and Product Evaluation. Kluwer, Dordrecht.

Grabe, E., Post, B., Nolan, F., and Farrar, K. (2000). Pitch accent realisation in four varieties of British English. Journal of Phonetics, 28, 161–186.

Grabe, E. (2001). The ivie labelling guide..

Gregory, M. and Altun, Y. (2004). Using conditional random fields to predict pitch accents in conversational speech. In ACL-04.

Grosjean, F., Grosjean, L., and Lane, H. (1979). The patterns of silence: Performance structures in sentence production. Cognitive Psychology, 11, 58–81.

Hamza, W., Bakis, R., Eide, E. M., Picheny, M. A., and Pitrelli, J. F. (2004). The IBM expressive speech synthesis system. In ICSLP-04, Jeju, Korea.

Harris, C. M. (1953). A study of the building blocks in speech. Journal of the Acoustical Society of America, 25(5), 962–969.

Hirschberg, J. (1993). Pitch Accent in Context: Predicting International Prominence from Text. Artificial Intelligence, 63(1-2), 305–340.

Hirschberg, J. and Pierrehumbert, J. B. (1986). The intonational structuring of discourse. In ACL-86, New York, pp. 136–144.

House, A. S., Williams, C. E., Hecker, M. H. L., and Kryter, K. D. (1965). Articulation-Testing Methods: Consonantal Differentiation with a Closed-Response Set. Journal of the Acoustical Society of America, 37, 158–166.

Huang, X., Acero, A., and Hon, H.-W. (2001). Spoken Language Processing: A Guide to Theory, Algorithm, and System Development. Prentice Hall, Upper Saddle River, NJ.

Hunt, A. J. and Black, A. W. (1996a). Unit selection in a concatenative speech synthesis system using a large speech database. In IEEE ICASSP-96, Atlanta, GA, Vol. 1, pp. 373–376. IEEE.

Hunt, A. J. and Black, A. W. (1996b). Unit selection in a concatenative speech synthesis system using a large speech database. In IEEE ICASSP-06, Vol. 1, pp. 373–376.

Jilka, M., Mohler, G., and Dogil, G. (1999). Rules for the generation of ToBI-based American English intonation. Speech Communication, 28(2), 83–108.

Jun, S.-A. (Ed.). (2005). Prosodic Typology and Transcription: A Unified Approach. Oxford University Press.

Klatt, D. H. (1975). Voice onset time, friction, and aspiration in word-initial consonant clusters. Journal of Speech and Hearing Research, 18, 686–706.

Klatt, D. H. (1982). The Klattalk text-to-speech conversion system. In IEEE ICASSP-82, pp. 1589–1592.

Klatt, D. H. (1979). Synthesis by rule of segmental durations in English sentences. In Lindblom, B. E. F. and Öhman, S. (Eds.), Frontiers of Speech Communication Research, pp. 287–299. Academic.

原书第 314 页

Klimt, B. and Yang, Y. (2004). The Enron corpus: A new dataset for email classification research. In Proceedings of the European Conference on Machine Learning, pp. 217–226. Springer.

Koehn, P., Abney, S. P., Hirschberg, J., and Collins, M. (2000). Improving intonational phrasing with syntactic information. In IEEE ICASSP-00.

Ladd, D. R. (1996). Intonational Phonology. Cambridge Studies in Linguistics. Cambridge University Press.

Langendoen, D. T. (1975). Finite-state parsing of phrase-structure languages and the status of readjustment rules in the grammar. Linguistic Inquiry, 6(4), 533–554.

Lawrence, W. (1953). The synthesis of speech from signals which have a low information rate.. In Jackson, W. (Ed.), Communication Theory, pp. 460–469. Butterworth.

Lee, S., Bresch, E., Adams, J., Kazemzadeh, A., and Narayanan, S. S. (2006). A study of emotional speech articulation using a fast magnetic resonance imaging technique. In ICSLP-06.

Liberman, M. Y. and Church, K. W. (1992). Text analysis and word pronunciation in text-to-speech synthesis. In Furui, S. and Sondhi, M. M. (Eds.), Advances in Speech Signal Processing, pp. 791–832. Marcel Dekker.

Liberman, M. Y. and Prince, A. (1977). On stress and linguistic rhythm. Linguistic Inquiry, 8, 249–336.

Liberman, M. Y. and Sproat, R. (1992). The stress and structure of modified noun phrases in English. In Sag, I. A. and Szabolcsi, A. (Eds.), Lexical Matters, pp. 131–181. CSLI, Stanford University.

Lucassen, J. and Mercer, R. L. (1984). An information theoretic approach to the automatic determination of phonemic baseforms. In IEEE ICASSP-84, Vol. 9, pp. 304–307.

Luk, R. W. P. and Damper, R. I. (1996). Stochastic phonographic transduction for english. Computer Speech and Language, 10(2), 133–153.

Marchand, Y. and Damper, R. I. (2000). A multi-strategy approach to improving pronunciation by analogy. Computational Linguistics, 26(2), 195–219.

Nakajima, S. and Hamada, H. (1988). Automatic generation of synthesis units based on context oriented clustering. In IEEE ICASSP-88, pp. 659–662.

Narayanan, S. S. and Alwan, A. (Eds.). (2004). Text to Speech Synthesis: New paradigms and advances. Prentice Hall.

Nenkova, A., Brenier, J., Kothari, A., Calhoun, S., Whitton, L., Beaver, D., and Jurafsky, D. (2007). To memorize or to predict: Prominence labeling in conversational speech. In NAACL-HLT 07.

Nespor, M. and Vogel, I. (1986). Prosodic phonology. Foris, Dordrecht.

Olive, J. and Liberman, M. (1979). A set of concatenative units for speech synthesis. Journal of the Acoustical Society of America, 65, S130.

Olive, J. P. (1977). Rule synthesis of speech from dyadic units. In ICASSP77, pp. 568–570.

Olive, J. P., van Santen, J. P. H., Möbius, B., and Shih, C. (1998). Synthesis. In Sproat, R. (Ed.), Multilingual Text-To-Speech Synthesis: The Bell Labs Approach, pp. 191–228. Kluwer, Dordrecht.

Ostendorf, M. and Veilleux, N. (1994). A hierarchical stochastic model for automatic prediction of prosodic boundary location. Computational Linguistics, 20(1).

Pan, S. and Hirschberg, J. (2000). Modeling local context for pitch accent prediction. In ACL-00, Hong Kong, pp. 233–240.

Pan, S. and McKeown, K. R. (1999). Word informativeness and automatic pitch accent modeling. In EMNLP/VLC-99.

Peterson, G. E., Wang, W. W.-Y., and Sivertsen, E. (1958). Segmentation techniques in speech synthesis. Journal of the Acoustical Society of America, 30(8), 739–742.

Pierrehumbert, J. B. (1980). The Phonology and Phonetics of English Intonation. Ph.D. thesis, MIT.

Pitrelli, J. F., Beckman, M. E., and Hirschberg, J. (1994). Evaluation of prosodic transcription labeling reliability in the ToBI framework. In ICSLP-94, Vol. 1, pp. 123–126.

Price, P. J., Ostendorf, M., Shattuck-Hufnagel, S., and Fong, C. (1991). The use of prosody in syntactic disambiguation. Journal of the Acoustical Society of America, 90(6).

Riley, M. D. (1992). Tree-based modelling for speech synthesis. In Bailly, G. and Beniot, C. (Eds.), Talking Machines: Theories, Models and Designs. North Holland, Amsterdam.

Sagisaka, Y. (1988). Speech synthesis by rule using an optimal selection of non-uniform synthesis units. In IEEE ICASSP-88, pp. 679–682.

Sagisaka, Y., Kaiki, N., Iwahashi, N., and Mimura, K. (1992). Atr – v-talk speech synthesis system. In ICSLP-92, Banff, Canada, pp. 483–486.

Sagisaka, Y., Campbell, N., and Higuchi, N. (Eds.). (1997). Computing Prosody: Computational Models for Processing Spontaneous Speech. Springer.

Schroder, M. (2006). Expressing degree of activation in synthetic speech. IEEE Transactions on Audio, Speech, and Language Processing, 14(4), 1128–1136.

Sejnowski, T. J. and Rosenberg, C. R. (1987). Parallel networks that learn to pronounce English text. Complex Systems, 1(1), 145–168.

Selkirk, E. (1986). On derived domains in sentence phonology. Phonology Yearbook, 3, 371–405.

Silverman, K., Beckman, M. E., Pitrelli, J. F., Ostendorf, M., Wightman, C., Price, P. J., Pierrehumbert, J. B., and Hirschberg, J. (1992). ToBI: a standard for labelling English prosody. In ICSLP-92, Vol. 2, pp. 867–870.

Spiegel, M. F. (2002). Proper name pronunciations for speech technology applications. In Proceedings of IEEE Workshop on Speech Synthesis, pp. 175–178.

Spiegel, M. F. (2003). Proper name pronunciations for speech technology applications. International Journal of Speech Technology, 6(4), 419–427.

原书第 315 页

Sproat, R. (1994). English noun-phrase prediction for text-to-speech. Computer Speech and Language, 8, 79–94.

Sproat, R. (1998a). Further issues in text analysis. In Sproat, R. (Ed.), Multilingual Text-To-Speech Synthesis: The Bell Labs Approach, pp. 89–114. Kluwer, Dordrecht.

Sproat, R. (Ed.). (1998b). Multilingual Text-To-Speech Synthesis: The Bell Labs Approach. Kluwer, Dordrecht.

Sproat, R., Black, A. W., Chen, S. F., Kumar, S., Ostendorf, M., and Richards, C. (2001). Normalization of non-standard words. Computer Speech & Language, 15(3), 287–333.

Steedman, M. (2003). Information-structural semantics for English intonation.

Stevens, K. N., Kasowski, S., and Fant, G. M. (1953). An electrical analog of the vocal tract. Journal of the Acoustical Society of America, 25(4), 734–742.

Streeter, L. (1978). Acoustic determinants of phrase boundary perception. Journal of the Acoustical Society of America, 63, 1582–1592.

Syrdal, A. K. and Conkie, A. (2004). Data-driven perceptually based join costs. In Proceedings of Fifth ISCA Speech Synthesis Workshop.

Syrdal, A. K., Wightman, C. W., Conkie, A., Stylianou, Y., Beutnagel, M., Schroeter, J., Strom, V., and Lee, K.-S. (2000). Corpus-based techniques in the AT&T NEXTGEN synthesis system. In ICSLP-00, Beijing.

Taylor, P. (2000). Analysis and synthesis of intonation using the Tilt model. Journal of the Acoustical Society of America, 107(3), 1697–1714.

Taylor, P. (2005). Hidden Markov Models for grapheme to phoneme conversion. In INTERSPEECH-05, Lisbon, Portugal, pp. 1973–1976.

Taylor, P. (2008). Text-to-speech synthesis. Manuscript.

Taylor, P. and Black, A. W. (1998). Assigning phrase breaks from part of speech sequences. Computer Speech and Language, 12, 99–117.

Taylor, P. and Isard, S. (1991). Automatic diphone segmentation. In EUROSPEECH-91, Genova, Italy.

Ternishi, R. and Umeda, N. (1968). Use of pronouncing dictionary in speech synthesis experiments. In 6th International Congress on Acoustics, Tokyo, Japan, pp. B155–158. †.

Tokuda, K., Kobayashi, T., and Imai, S. (1995a). Speech parameter generation from hmm using dynamic features. In IEEE ICASSP-95.

Tokuda, K., Masuko, T., and Yamada, T. (1995b). An algorithm for speech parameter generation from continuous mixture hmms with dynamic features. In EUROSPEECH-95, Madrid.

Tokuda, K., Yoshimura, T., Masuko, T., Kobayashi, T., and Kitamura, T. (2000). Speech parameter generation algorithms for hmm-based speech synthesis. In IEEE ICASSP-00.

Tokuda, K., Zen, H., and Kitamura, T. (2003). Trajectory modeling based on hmms with the explicit relationship between static and dynamic features. In EUROSPEECH-03.

Umeda, N., Matui, E., Suzuki, T., and Omura, H. (1968). Synthesis of fairy tale using an analog vocal tract. In 6th International Congress on Acoustics, Tokyo, Japan, pp. B159–162.

Umeda, N. (1976). Linguistic rules for text-to-speech synthesis. Proceedings of the IEEE, 64(4), 443–451.

van Santen, J. P. H. (1994). Assignment of segmental duration in text-to-speech synthesis. Computer Speech and Language, 8(95–128).

van Santen, J. P. H. (1997). Segmental duration and speech timing. In Sagisaka, Y., Campbell, N., and Higuchi, N. (Eds.), Computing Prosody: Computational Models for Processing Spontaneous Speech. Springer.

van Santen, J. P. H. (1998). Timing. In Sproat, R. (Ed.), Multilingual Text-To-Speech Synthesis: The Bell Labs Approach, pp. 115–140. Kluwer, Dordrecht.

van Santen, J. P. H., Sproat, R., Olive, J. P., and Hirschberg, J. (Eds.). (1997). Progress in Speech Synthesis. Springer.

Veldhuis, R. (2000). Consistent pitch marking. In ICSLP-00, Beijing, China.

Venditti, J. J. (2005). The j_tobi model of japanese intonation. In Jun, S.-A. (Ed.), Prosodic Typology and Transcription: A Unified Approach. Oxford University Press.

Voiers, W., Sharpley, A., and Hehmsoth, C. (1975). Research on diagnostic evaluation of speech intelligibility. Research Report AFCRL-72-0694.

Wang, M. Q. and Hirschberg, J. (1992). Automatic classification of intonational phrasing boundaries. Computer Speech and Language, 6(2), 175–196.

Wouters, J. and Macon, M. (1998). Perceptual evaluation of distance measures for concatenative speech synthesis. In ICSLP-98, Sydney, pp. 2747–2750.

Yarowsky, D. (1997). Homograph disambiguation in text-to-speech synthesis. In van Santen, J. P. H., Sproat, R., Olive, J. P., and Hirschberg, J. (Eds.), Progress in Speech Synthesis, pp. 157–172. Springer.

Yuan, J., Brenier, J. M., and Jurafsky, D. (2005). Pitch accent prediction: Effects of genre and speaker. In EUROSPEECH05.

原书第 316 页

9

AUTOMATIC SPEECH RECOGNITION

When Frederic was a little lad he proved so brave and daring,

His father thought he'd 'prentice him to some career seafaring.

I was, alas! his nurs'rymaid, and so it fell to my lot

To take and bind the promising boy apprentice to a pilot —

A life not bad for a hardy lad, though surely not a high lot,

Though I'm a nurse, you might do worse than make your boy a pilot.

I was a stupid nurs'rymaid, on breakers always steering,

And I did not catch the word aright, through being hard of hearing;

Mistaking my instructions, which within my brain did gyrate,

I took and bound this promising boy apprentice to a pirate.

The Pirates of Penzance, Gilbert and Sullivan, 1877

Alas, this mistake by nurserymaid Ruth led to Frederic's long indenture as a pirate and, due to a slight complication involving 21st birthdays and leap years, nearly led to 63 extra years of apprenticeship. The mistake was quite natural, in a Gilbert-and-Sullivan sort of way; as Ruth later noted, "The two words were so much alike!" True, true; spoken language understanding is a difficult task, and it is remarkable that humans do as well at it as we do. The goal of automatic speech recognition (ASR) research is to address this problem computationally by building systems that map from an acoustic signal to a string of words. Automatic speech understanding (ASU) extends this goal to producing some sort of understanding of the sentence, rather than just the words.

The general problem of automatic transcription of speech by any speaker in any environment is still far from solved. But recent years have seen ASR technology mature to the point where it is viable in certain limited domains. One major application area is in human-computer interaction. While many tasks are better solved with visual or pointing interfaces, speech has the potential to be a better interface than the keyboard for tasks where full natural language communication is useful, or for which keyboards are not appropriate. This includes hands-busy or eyes-busy applications, such as where the user has objects to manipulate or equipment to control. Another important application area is telephony, where speech recognition is already used for example in spoken dialogue systems for entering digits, recognizing "yes" to accept collect calls, finding out airplane or train information, and call-routing ("Accounting, please", "Prof. Regier, please"). In some applications, a multimodal interface combining speech and pointing can be more efficient than a graphical user interface without speech (Cohen et al., 1998). Finally, ASR is applied to dictation, that is, transcription of extended

原书第 317 页

monologue by a single specific speaker. Dictation is common in fields such as law and is also important as part of augmentative communication (interaction between computers and humans with some disability resulting in the inability to type, or the inability to speak). The blind Milton famously dictated Paradise Lost to his daughters, and Henry James dictated his later novels after a repetitive stress iniurv.

Before turning to architectural details, let's discuss some of the parameters and the state of the art of the speech recognition task. One dimension of variation in speech recognition tasks is the vocabulary size. Speech recognition is easier if the number of distinct words we need to recognize is smaller. So tasks with a two word vocabulary, like yes versus no detection, or an eleven word vocabulary, like recognizing sequences of digits, in what is called the digits task task, are relatively easy. On the other hand, tasks with large vocabularies, like transcribing human-human telephone conversations, or transcribing broadcast news, tasks with vocabularies of 64,000 words or more, are much harder.

A second dimension of variation is how fluent, natural, or conversational the speech is. Isolated word recognition, in which each word is surrounded by some sort of pause, is much easier than recognizing continuous speech, in which words run into each other and have to be segmented. Continuous speech tasks themselves vary greatly in difficulty. For example, human-to-machine speech turns out to be far easier to recognize than human-to-human speech. That is, recognizing speech of humans talking to machines, either reading out loud in read speech (which simulates the dictation task), or conversing with speech dialogue systems, is relatively easy. Recognizing the speech of two humans talking to each other, in conversational speech recognition, for example for transcribing a business meeting or a telephone conversation, is much harder. It seems that when humans talk to machines, they simplify their speech quite a bit, talking more slowly and more clearly.

A third dimension of variation is channel and noise. Commercial dictation systems, and much laboratory research in speech recognition, is done with high quality, head mounted microphones. Head mounted microphones eliminate the distortion that occurs in a table microphone as the speakers head moves around. Noise of any kind also makes recognition harder. Thus recognizing a speaker dictating in a quiet office is much easier than recognizing a speaker dictating in a noisy car on the highway with the window open.

A final dimension of variation is accent or speaker-class characteristics. Speech is easier to recognize if the speaker is speaking a standard dialect, or in general one that matches the data the system was trained on. Recognition is thus harder on foreign-accented speech, or speech of children (unless the system was specifically trained on exactly these kinds of speech).

Table 9.1 shows the rough percentage of incorrect words (the word error rate, or WER, defined on page 45) from state-of-the-art systems on a range of different ASR tasks.

Variation due to noise and accent increases the error rates quite a bit. The word error rate on strongly Japanese-accented or Spanish accented English has been reported to be about 3 to 4 times higher than for native speakers on the same task (Tomokiyo, 2001). And adding automobile noise with a 10dB SNR (signal-to-noise ratio) can cause error rates to go up by 2 to 4 times.

原书第 318 页
TaskVocabularyError Rate %
TI Digits11 (zero-nine, oh).5
Wall Street Journal read speech5,0003
Wall Street Journal read speech20,0003
Broadcast News64,000+10
Conversational Telephone Speech (CTS)64,000+20
Figure 9.1 Rough word error rates (% of words misrecognized) reported around 2006 for ASR on various tasks; the error rates for Broadcast News and CTS are based on particular training and test scenarios and should be taken as ballpark numbers; error rates for differently defined tasks may range up to a factor of two.

In general, these error rates go down every year, as speech recognition performance has improved quite steadily. One estimate is that performance has improved roughly 10 percent a year over the last decade (Deng and Huang, 2004), due to a combination of algorithmic improvements and Moore's law.

While the algorithms we describe in this chapter are applicable across a wide variety of these speech tasks, we chose to focus this chapter on the fundamentals of one crucial area: Large-Vocabulary Continuous Speech Recognition (LVCSR). Large-vocabulary generally means that the systems have a vocabulary of roughly 20,000 to 60,000 words. We saw above that $ \underline{\text{continuous}} $ means that the words are run together naturally. Furthermore, the algorithms we will discuss are generally $ \underline{\text{speaker-independent}} $; that is, they are able to recognize speech from people whose speech the system has never been exposed to before.

The dominant paradigm for LVCSR is the HMM, and we will focus on this approach in this chapter. Previous chapters have introduced most of the core algorithms used in HMM-based speech recognition. Ch. 7 introduced the key phonetic and phonological notions of phone, syllable, and intonation. Ch. 5 and Ch. 6 introduced the use of the Bayes rule, the Hidden Markov Model (HMM), the Viterbi algorithm, and the Baum-Welch training algorithm. Ch. 4 introduced the N-gram language model and the perplexity metric. In this chapter we begin with an overview of the architecture for HMM speech recognition, offer an all-too-brief overview of signal processing for feature extraction and the extraction of the important MFCC features, and then introduce Gaussian acoustic models. We then continue with how Viterbi decoding works in the ASR context, and give a complete summary of the training procedure for ASR, called embedded training. Finally, we introduce word error rate, the standard evaluation metric. The next chapter will continue with some advanced ASR topics.

9.1 SPEECH RECOGNITION ARCHITECTURE

The task of speech recognition is to take as input an acoustic waveform and produce as output a string of words. HMM-based speech recognition systems view this task using the metaphor of the noisy channel. The intuition of the noisy channel model (see Fig. 9.2) is to treat the acoustic waveform as an “noisy” version of the string of words, i.e., a version that has been passed through a noisy communications channel.

原书第 319 页

This channel introduces “noise” which makes it hard to recognize the “true” string of words. Our goal is then to build a model of the channel so that we can figure out how it modified this “true” sentence and hence recover it.

The insight of the noisy channel model is that if we know how the channel distorts the source, we could find the correct source sentence for a waveform by taking every possible sentence in the language, running each sentence through our noisy channel model, and seeing if it matches the output. We then select the best matching source sentence as our desired source sentence.

Image
Figure 9.2 The noisy channel model. We search through a huge space of potential “source” sentences and choose the one which has the highest probability of generating the “noisy” sentence. We need models of the prior probability of a source sentence (N-grams), the probability of words being realized as certain strings of phones (HMM lexicons), and the probability of phones being realized as acoustic or spectral features (Gaussian Mixture Models).

Implementing the noisy-channel model as we have expressed it in Fig. 9.2 requires solutions to two problems. First, in order to pick the sentence that best matches the noisy input we will need a complete metric for a “best match”. Because speech is so variable, an acoustic input sentence will never exactly match any model we have for this sentence. As we have suggested in previous chapters, we will use probability as our metric. This makes the speech recognition problem a special case of Bayesian inference, a method known since the work of Bayes (1763). Bayesian inference or Bayesian classification was applied successfully by the 1950s to language problems like optical character recognition (Bledsoe and Browning, 1959) and to authorship attribution tasks like the seminal work of Mosteller and Wallace (1964) on determining the authorship of the Federalist papers. Our goal will be to combine various probabilistic models to get a complete estimate for the probability of a noisy acoustic observation-sequence given a candidate source sentence. We can then search through the space of all sentences, and choose the source sentence with the highest probability.

Second, since the set of all English sentences is huge, we need an efficient algorithm that will not search through all possible sentences, but only ones that have a good

原书第 320 页

chance of matching the input. This is the decoding or search problem, which we have already explored with the Viterbi decoding algorithm for HMMs in Ch. 5 and Ch. 6. Since the search space is so large in speech recognition, efficient search is an important part of the task, and we will focus on a number of areas in search.

In the rest of this introduction we will review the probabilistic or Bayesian model for speech recognition that we introduced for part-of-speech tagging in Ch. 5. We then introduce the various components of a modern HMM-based ASR system.

Recall that the goal of the probabilistic noisy channel architecture for speech recognition can be summarized as follows:

“What is the most likely sentence out of all sentences in the language L given some acoustic input O?”

We can treat the acoustic input $O$ as a sequence of individual “symbols” or “observations” (for example by slicing up the input every 10 milliseconds, and representing each slice by floating-point values of the energy or frequencies of that slice). Each index then represents some time interval, and successive $o_i$ indicate temporally consecutive slices of the input (note that capital letters will stand for sequences of symbols and lower-case letters for individual symbols):

$$ O=o_{1},o_{2},o_{3},\ldots,o_{t} $$

Similarly, we treat a sentence as if it were composed of a string of words:

$$ \boldsymbol{W}=\boldsymbol{w}_{1},\boldsymbol{w}_{2},\boldsymbol{w}_{3},\cdots,\boldsymbol{w}_{n} $$

Both of these are simplifying assumptions; for example dividing sentences into words is sometimes too fine a division (we'd like to model facts about groups of words rather than individual words) and sometimes too gross a division (we need to deal with morphology). Usually in speech recognition a word is defined by orthography (after mapping every word to lower-case): oak is treated as a different word than oaks, but the auxiliary can (“can you tell me…”) is treated as the same word as the noun can (“I need a can of…”).

The probabilistic implementation of our intuition above, then, can be expressed as follows:

$$ \hat{W}=\underset{W\in\mathcal{L}}{\mathrm{a r g m a x}}P(W|O) $$

Recall that the function $ \argmax_{x} f(x) $ means “the x such that $ f(x) $ is largest”. Equation (9.3) is guaranteed to give us the optimal sentence W; we now need to make the equation operational. That is, for a given sentence W and acoustic sequence O we need to compute $ P(W|O) $. Recall that given any probability $ P(x|y) $, we can use Bayes’ rule to break it down as follows:

$$ P(x|y)=\frac{P(y|x)P(x)}{P(y)} $$

We saw in Ch. 5 that we can substitute (9.4) into (9.3) as follows:

原书第 321 页

$$ \hat{W}=\underset{W\in\mathcal{L}}{\mathrm{a r g m a x}}\frac{P(O|W)P(W)}{P(O)} $$

The probabilities on the right-hand side of (9.5) are for the most part easier to compute than $ P(W|O) $. For example, $ P(W) $, the prior probability of the word string itself is exactly what is estimated by the n-gram language models of Ch. 4. And we will see below that $ P(O|W) $ turns out to be easy to estimate as well. But $ P(O) $, the probability of the acoustic observation sequence, turns out to be harder to estimate. Luckily, we can ignore $ P(O) $ just as we saw in Ch. 5. Why? Since we are maximizing over all possible sentences, we will be computing $ \frac{P(O|W)P(W)}{P(O)} $ for each sentence in the language. But $ P(O) $ doesn't change for each sentence! For each potential sentence we are still examining the same observations O, which must have the same probability $ P(O) $. Thus:

$$ \hat{W}=\underset{W\in\mathcal{L}}{\operatorname{a r g m a x}}\frac{P(O|W)P(W)}{P(O)}=\underset{W\in\mathcal{L}}{\operatorname{a r g m a x}}P(O|W)P(W) $$

To summarize, the most probable sentence $ W $ given some observation sequence $ O $ can be computed by taking the product of two probabilities for each sentence, and choosing the sentence for which this product is greatest. The general components of the speech recognizer which compute these two terms have names; $ P(W) $, the prior probability, is computed by the language model. while $ P(O|W) $, the observation likelihood, is computed by the acoustic model.

$$ \underset{(w\in\mathcal{L})}{\hat{W}=\operatornamewithlimits{a r g m a x}_{W\in\mathcal{L}}\overbrace{P(O|W)}^{l i k e l i h o o d}\overbrace{P(W)}^{p r i o r} $$

The language model (LM) prior $ P(W) $ expresses how likely a given string of words is to be a source sentence of English. We have already seen in Ch. 4 how to compute such a language model prior $ P(W) $ by using N-gram grammars. Recall that an N-gram grammar lets us assign a probability to a sentence by computing:

$$ P(w_{1}^{n})\approx\prod_{k=1}^{n}P(w_{k}|w_{k-N+1}^{k-1}) $$

This chapter will show how the HMM we covered in Ch. 6 can be used to build an Acoustic Model (AM) which computes the likelihood $ P(O|W) $. Given the AM and LM probabilities, the probabilistic model can be operationalized in a search algorithm so as to compute the maximum probability word string for a given acoustic waveform. Fig. 9.3 shows the components of an HMM speech recognizer as it processes a single utterance, indicating the computation of the prior and likelihood. The figure shows the recognition process in three stages. In the feature extraction or signal processing stage, the acoustic waveform is sampled into frames (usually of 10, 15, or 20 milliseconds) which are transformed into spectral features. Each time window is thus represented by a vector of around 39 features representing this spectral information as well as information about energy and spectral change. Sec. 9.3 gives an (unfortunately brief) overview of the feature extraction process.

原书第 322 页

In the acoustic modeling or phone recognition stage, we compute the likelihood of the observed spectral feature vectors given linguistic units (words, phones, subparts of phones). For example, we use Gaussian Mixture Model (GMM) classifiers to compute for each HMM state $q$, corresponding to a phone or subphone, the likelihood of a given feature vector given this phone $p(o|q)$. A (simplified) way of thinking of the output of this stage is as a sequence of probability vectors, one for each time frame, each vector at each time frame containing the likelihoods that each phone or subphone unit generated the acoustic feature vector observation at that time.

Finally, in the decoding phase, we take the acoustic model (AM), which consists of this sequence of acoustic likelihoods, plus an HMM dictionary of word pronunciations, combined with the language model (LM) (generally an N-gram grammar), and output the most likely sequence of words. An HMM dictionary, as we will see in Sec. 9.2, is a list of word pronunciations, each pronunciation represented by a string of phones. Each word can then be thought of as an HMM, where the phones (or sometimes subphones) are states in the HMM, and the Gaussian likelihood estimators supply the HMM output likelihood function for each state. Most ASR systems use the Viterbi algorithm for decoding, speeding up the decoding with wide variety of sophisticated augmentations such as pruning, fast-match, and tree-structured lexicons.

Image
Figure 9.3 Schematic architecture for a (simplified) speech recognizer decoding a single sentence. A real recognizer is more complex since various kinds of pruning and fast matches are needed for efficiency. This architecture is only for decoding; we also need a separate architecture for training parameters.
原书第 323 页

9.2 APPLYING THE HIDDEN MARKOV MODEL TO SPEECH

Let's turn now to how the HMM model is applied to speech recognition. We saw in Ch. 6 that a Hidden Markov Model is characterized by the following components:

$ Q = q_1 q_2 \ldots q_N $ a set of states

$ A = a_{01} a_{02} \ldots a_{n1} \ldots a_{nn} $ a transition probability matrix A, each $ a_{ij} $ representing the probability of moving from state i to state j, s.t. $ \sum_{j=1}^{n} a_{ij} = 1 \quad \forall i $

$ O = o_1 o_2 \ldots o_N $ a set of observations, each one drawn from a vocabulary $ V = \nu_1, \nu_2, \ldots, \nu_V $.

$ B = b_i(o_t) $ A set of observation likelihoods:, also called emission probabilities, each expressing the probability of an observation $ o_t $ being generated from a state i.

$ q_0, q_{end} $ a special start and end state which are not associated with observations.

Furthermore, the chapter introduced the Viterbi algorithm for decoding HMMs, and the Baum-Welch or Forward-Backward algorithm for training HMMs.

All of these facets of the HMM paradigm play a crucial role in ASR. We begin here by discussing how the states, transitions, and observations map into the speech recognition task. We will return to the ASR applications of Viterbi decoding in Sec. 9.6. The extensions to the Baum-Welch algorithms needed to deal with spoken language are covered in Sec. 9.4 and Sec. 9.7.

Recall the examples of HMMs we saw earlier in the book. In Ch. 5, the hidden states of the HMM were parts-of-speech, the observations were words, and the HMM decoding task mapped a sequence of words to a sequence of parts-of-speech. In Ch. 6, the hidden states of the HMM were weather, the observations were 'ice-cream consumptions', and the decoding task was to determine the weather sequence from a sequence of ice-cream consumption. For speech, the hidden states are phones, parts of phones, or words, each observation is information about the spectrum and energy of the waveform at a point in time, and the decoding process maps this sequence of acoustic information to phones and words.

The observation sequence for speech recognition is a sequence of acoustic feature vectors. Each acoustic feature vector represents information such as the amount of energy in different frequency bands at a particular point in time. We will return in Sec. 9.3 to the nature of these observations, but for now we'll simply note that each observation consists of a vector of 39 real-valued features indicating spectral information. Observations are generally drawn every 10 milliseconds, so 1 second of speech requires 100 spectral feature vectors, each vector of length 39.

The hidden states of Hidden Markov Models can be used to model speech in a number of different ways. For small tasks, like digit recognition, (the recognition of the 10 digit words zero through nine), or for yes-no recognition (recognition of the two words yes and no), we could build an HMM whose states correspond to entire words.

原书第 324 页

For most larger tasks, however, the hidden states of the HMM correspond to phone-like units, and words are sequences of these phone-like units.

Let's begin by describing an HMM model in which each state of an HMM corresponds to a single phone (if you've forgotten what a phone is, go back and look again at the definition in Ch. 7). In such a model, a word HMM thus consists of a sequence of HMM states concatenated together. Fig. 9.4 shows a schematic of the structure of a basic phone-state HMM for the word six.

Image
Figure 9.4 An HMM for the word six, consisting of four emitting states and two non-emitting states, the transition probabilities A, the observation probabilities B, and a sample observation sequence.

Note that only certain connections between phones exist in Fig. 9.4. In the HMMs described in Ch. 6, there were arbitrary transitions between states; any state could transition to any other. This was also in principle true of the HMMs for part-of-speech tagging in Ch. 5; although the probability of some tag transitions was low, any tag could in principle follow any other tag. Unlike in these other HMM applications, HMM models for speech recognition usually do not allow arbitrary transitions. Instead, they place strong constraints on transitions based on the sequential nature of speech. Except in unusual cases, HMMs for speech don't allow transitions from states to go to earlier states in the word; in other words, states can transition to themselves or to successive states. As we saw in Ch. 6, this kind of left-to-right HMM structure is called a Bakis network.

The most common model used for speech, illustrated in a simplified form in Fig. 9.4 is even more constrained, allowing a state to transition only to itself (self-loop) or to a single succeeding state. The use of self-loops allows a single phone to repeat so as to cover a variable amount of the acoustic input. Phone durations vary hugely, dependent on the phone identity, the speaker's rate of speech, the phonetic context, and the level of prosodic prominence of the word. Looking at the Switchboard corpus, the phone [aa] varies in length from 7 to 387 milliseconds (1 to 40 frames), while the phone [z] varies in duration from 7 milliseconds to more than 1.3 seconds (130 frames) in some utterances! Self-loops thus allow a single state to be repeated many times.

For very simple speech tasks (recognizing small numbers of words such as the 10 digits), using an HMM state to represent a phone is sufficient. In general LVCSR tasks, however, a more fine-grained representation is necessary. This is because phones can last over 1 second, i.e., over 100 frames, but the 100 frames are not acoustically identical. The spectral characteristics of a phone, and the amount of energy, vary dramatically across a phone. For example, recall from Ch. 7 that stop consonants have a closure portion, which has very little acoustic energy, followed by a release burst. Similarly, diphthongs are vowels whose F1 and F2 change significantly. Fig. 9.5 shows

原书第 325 页

these large changes in spectral characteristics over time for each of the two phones in the word "Ike", ARPAbet [ay k].

Image
Figure 9.5 The two phones of the word "Ike", pronounced [ay k]. Note the continuous changes in the [ay] vowel on the left, as F2 rises and F1 falls, and the sharp differences between the silence and release parts of the [k] stop.

To capture this fact about the non-homogeneous nature of phones over time, in LVCSR we generally model a phone with more than one HMM state. The most common configuration is to use three HMM states, a beginning, middle, and end state. Each phone thus consists of 3 emitting HMM states instead of one (plus two non-emitting states at either end), as shown in Fig. 9.6. It is common to reserve the word model or phone model to refer to the entire 5-state phone HMM, and use the word HMM state (or just state for short) to refer to each of the 3 individual subphones HMM states.

Image
Figure 9.6 A standard 5-state HMM model for a phone, consisting of three emitting states (corresponding to the transition-in, steady state, and transition-out regions of the phone) and two non-emitting states.

To build a HMM for an entire word using these more complex phone models, we can simply replace each phone of the word model in Fig. 9.4 with a 3-state phone HMM. We replace the non-emitting start and end states for each phone model with

原书第 326 页

transitions directly to the emitting state of the preceding and following phone, leaving only two non-emitting states for the entire word. Fig. 9.7 shows the expanded word.

Image
Figure 9.7 A composite word model for “six”, [s ih k s], formed by concatenating four phone models, each with three emitting states.

In summary, an HMM model of speech recognition is parameterized by:

$$ Q=q_{1}q_{2}\ldots q_{N} $$

a set of states corresponding to subphones

$$ A=a_{01}a_{02}\ldots a_{n1}\ldots a_{nn} $$

a transition probability matrix $ A $, each $ a_{ij} $ representing the probability for each subphone of taking a self-loop or going to the next subphone.

$$ \boldsymbol{B}=\boldsymbol{b}_{i}(o_{t}) $$

A set of observation likelihoods:, also called emission probabilities, each expressing the probability of a cepstral feature vector (observation $ o_{t} $) being generated from subphone state i.

Another way of looking at the A probabilities and the states Q is that together they represent a lexicon: a set of pronunciations for words, each pronunciation consisting of a set of subphones, with the order of the subphones specified by the transition probabilities A.

We have now covered the basic structure of HMM states for representing phones and words in speech recognition. Later in this chapter we will see further augmentations of the HMM word model shown in Fig. 9.7, such as the use of triphone models which make use of phone context, and the use of special phones to model silence. First, though, we need to turn to the next component of HMMs for speech recognition: the observation likelihoods. And in order to discuss observation likelihoods, we first need to introduce the actual acoustic observations: feature vectors. After discussing these in Sec. 9.3, we turn in Sec. 9.4 the acoustic model and details of observation likelihood computation. We then re-introduce Viterbi decoding and show how the acoustic model and language model are combined to choose the best sentence.

9.3 FEATURE EXTRACTION: MFCC VECTORS

Our goal in this section is to describe how we transform the input waveform into a sequence of acoustic feature vectors, each vector representing the information in a small time window of the signal. While there are many possible such feature representations, by far the most common in speech recognition is the MFCC, the mel frequency cepstral coefficients. These are based on the important idea of the cepstrum. We will give a relatively high-level description of the process of extraction of MFCCs from a

原书第 327 页
Image
Figure 9.8 Extracting a sequence of 39-dimensional MFCC feature vectors from a quantized digitized waveform

waveform; we strongly encourage students interested in more detail to follow up with a speech signal processing course.

We begin by repeating from Sec. ?? the process of digitizing and quantizing an analog speech waveform. Recall that the first step in processing speech is to convert the analog representations (first air pressure, and then analog electric signals in a microphone), into a digital signal. This process of analog-to-digital conversion has two steps: sampling and quantization. A signal is sampled by measuring its amplitude at a particular time; the sampling rate is the number of samples taken per second. In order to accurately measure a wave, it is necessary to have at least two samples in each cycle: one measuring the positive part of the wave and one measuring the negative part. More than two samples per cycle increases the amplitude accuracy, but less than two samples will cause the frequency of the wave to be completely missed. Thus the maximum frequency wave that can be measured is one whose frequency is half the sample rate (since every cycle needs two samples). This maximum frequency for a given sampling rate is called the Nyquist frequency. Most information in human speech is in frequencies below 10,000 Hz; thus a 20,000 Hz sampling rate would be necessary for complete accuracy. But telephone speech is filtered by the switching network, and only frequencies less than 4,000 Hz are transmitted by telephones. Thus an 8,000 Hz sampling rate is sufficient for telephone-bandwidth speech like the Switchboard corpus. A 16,000 Hz sampling rate (sometimes called wideband) is often used for microphone speech.

Even an 8,000 Hz sampling rate requires 8000 amplitude measurements for each second of speech, and so it is important to store the amplitude measurement efficiently. They are usually stored as integers, either 8-bit (values from -128 to 127) or 16 bit (values from -3276 to 32767). This process of representing real-valued numbers as integers is called quantization because there is a minimum granularity (the quantum size) and all values which are closer together than this quantum size are represented identically.

We refer to each sample in the digitized quantized waveform as $ x[n] $, where n is an index over time. Now that we have a digitized, quantized representation of the waveform, we are ready to extract MFCC features. The seven steps of this process are shown in Fig. 9.8 and individually described in each of the following sections.

← 8.4.1 Building a diphone database9.3.1 Preemphasis →