← 学习库 Speech and Language Processing 本册目录

7.4.6 The Source-Filter Model

Why do different vowels have different spectral signatures? As we briefly mentioned above, the formants are caused by the resonant cavities of the mouth. The source-filter model is a way of explaining the acoustics of a sound by modeling how the pulses produced by the glottis (the source) are shaped by the vocal tract (the filter).

Let's see how this works. Whenever we have a wave such as the vibration in air caused by the glottal pulse, the wave also has harmonics. A harmonic is another wave whose frequency is a multiple of the fundamental wave. Thus for example a 115 Hz glottal fold vibration leads to harmonics (other waves) of 230 Hz, 345 Hz, 460 Hz, and

原书第 264 页

so on on. In general each of these waves will be weaker, i.e. have much less amplitude than the wave at the fundamental frequency.

It turns out, however, that the vocal tract acts as a kind of filter or amplifier; indeed any cavity such as a tube causes waves of certain frequencies to be amplified, and others to be damped. This amplification process is caused by the shape of the cavity; a given shape will cause sounds of a certain frequency to resonate and hence be amplified. Thus by changing the shape of the cavity we can cause different frequencies to be amplified.

Now when we produce particular vowels, we are essentially changing the shape of the vocal tract cavity by placing the tongue and the other articulators in particular positions. The result is that different vowels cause different harmonics to be amplified.

So a wave of the same fundamental frequency passed through different vocal tract positions will result in different harmonics being amplified.

We can see the result of this amplification by looking at the relationship between the shape of the oral tract and the corresponding spectrum. Fig. 7.26 shows the vocal tract position for three vowels and a typical resulting spectrum. The formants are places in the spectrum where the vocal tract happens to amplify particular harmonic frequencies.

Image
Figure 7.26 Visualizing the vocal tract position as a filter: the tongue positions for three English vowels and the resulting smoothed spectra showing F1 and F2. Tongue positions modeled after Ladefoged (1996).

7.5 PHONETIC RESOURCES

A wide variety of phonetic resources can be drawn on for computational work. One key set of resources are pronunciation dictionaries. Such on-line phonetic dictionaries give phonetic transcriptions for each word. Three commonly used on-line dictionaries for English are the CELEX, CMUdict, and PRONLEX lexicons; for other

原书第 265 页

languages, the LDC has released pronunciation dictionaries for Egyptian Arabic, German, Japanese, Korean, Mandarin, and Spanish. All these dictionaries can be used for both speech recognition and synthesis work.

The CELEX dictionary (Baayen et al., 1995) is the most richly annotated of the dictionaries. It includes all the words in the Oxford Advanced Learner's Dictionary (1974) (41,000 lemmata) and the Longman Dictionary of Contemporary English (1978) (53,000 lemmata), in total it has pronunciations for 160,595 wordforms. Its (British rather than American) pronunciations are transcribed using an ASCII version of the IPA called SAM. In addition to basic phonetic information like phone strings, syllabification, and stress level for each syllable, each word is also annotated with morphological, part of speech, syntactic, and frequency information. CELEX (as well as CMU and PRON-LEX) represent three levels of stress: primary stress, secondary stress, and no stress. For example, some of the CELEX information for the word dictionary includes multiple pronunciations ('dIk-S@n-rI and 'dIk-S@-n@-rI, corresponding to ARPABET [dih k sh ax n r ih] and [d ih k sh ax n ax r ih] respectively), together with the CVskelata for each one ([CVC][CVC][CV] and [CVC][CV][CV][CV]), the frequency of the word, the fact that it is a noun, and its morphological structure (diction+ary).

The free CMU Pronouncing Dictionary (CMU, 1993) has pronunciations for about 125,000 wordforms. It uses an 39-phone ARPabet-derived phoneme set. Transcriptions are phonemic, and thus instead of marking any kind of surface reduction like flapping or reduced vowels, it marks each vowel with the number 0 (unstressed) 1 (stressed), or 2 (secondary stress). Thus the word tiger is listed as [T AY1 G ER0] the word table as [T EY1 B AH0 L], and the word dictionary as [D IH1 K SH AH0 N EH2 R IY0]. The dictionary is not syllabified, although the nucleus is implicitly marked by the (numbered) vowel.

The PRONLEX dictionary (LDC, 1995) was designed for speech recognition and contains pronunciations for 90,694 wordforms. It covers all the words used in many years of the Wall Street Journal, as well as the Switchboard Corpus. PRONLEX has the advantage that it includes many proper names (20,000, where CELEX only has about 1000). Names are important for practical applications, and they are both frequent and difficult; we return to a discussion of deriving name pronunciations in Ch. 8.

Another useful resource is a phonetically annotated corpus, in which a collection of waveforms is hand-labeled with the corresponding string of phones. Two important phonetic corpora in English are the TIMIT corpus and the Switchboard Transcription Project corpus.

The TIMIT corpus (NIST, 1990) was collected as a joint project between Texas Instruments (TI), MIT, and SRI. It is a corpus of 6300 read sentences, where 10 sentences each from 630 speakers. The 6300 sentences were drawn from a set of 2342 pre-designed sentences, some selected to have particular dialect shibboleths, others to maximize phonetic diphone coverage. Each sentence in the corpus was phonetically hand-labeled, the sequence of phones was automatically aligned with the sentence wavefile, and then the automatic phone boundaries were manually hand-corrected (Seneff and Zue, 1988). The result is a time-aligned transcription; a transcription in which each phone in the transcript is associated with a start and end time in the waveform; we showed a graphical example of a time-aligned transcription in Fig. 7.17.

The phoneset for TIMIT, and for the Switchboard Transcription Project corpus be-

原书第 266 页

low, is a more detailed one than the minimal phonemic version of the ARPAbet. In particular, these phonetic transcriptions make use of the various reduced and rare phones mentioned in Fig. 7.1 and Fig. 7.2; the flap [dx], glottal stop [q], reduced vowels [ax], [ix], [axr], voiced allophone of [h] ([hv]), and separate phones for stop closure ([dcl], [tcl], etc) and release ([d], [t], etc). An example transcription is shown in Fig. 7.27.

shehadyourdarksuitingreasywashwaterallyear
sh iyhv ae dcljh axrdcl d aa r kcls ux qengcl g ri y s ixw aa shq w aa dx axr qaa ly ix axr
Figure 7.27 Phonetic transcription from the TIMIT corpus. Note palatalization of [d] in had, unreleased final stop in dark, glottalization of final [t] in suit to [q], and flap of [t] in water. The TIMIT corpus also includes time-alignments for each phone (not shown).

Where TIMIT is based on read speech, the more recent Switchboard Transcription Project corpus is based on the Switchboard corpus of conversational speech. This phonetically-annotated portion consists of approximately 3.5 hours of sentences extracted from various conversations (Greenberg et al., 1996). As with TIMIT, each annotated utterance contains a time-aligned transcription. The Switchboard transcripts are time-aligned at the syllable level rather than at the phone level; thus a transcript consists of a sequence of syllables with the start and end time of each syllable in the corresponding wavefile. Fig. 7.28 shows an example from the Switchboard Transcription Project, for the phrase they're kind of in between right now:

0.470 dh er0.640 k aa0.720 n ax0.900 v ih m0.953 b ix1.279 t w iy n1.410 r ay1.630 n aw
Figure 7.28 Phonetic transcription of the Switchboard phrase they're kind of in between right now. Note vowel reduction in they're and of, coda deletion in kind and right, and resyllabification (the [v] of of attaches as the onset of in). Time is given in number of seconds from beginning of sentence to start of each syllable.

Phonetically transcribed corpora are also available for other languages; the Kiel corpus of German is commonly used, as are various Mandarin corpora transcribed by the Chinese Academy of Social Sciences (Li et al., 2000).

In addition to resources like dictionaries and corpora, there are many useful phonetic software tools. One of the most versatile is the free Praat package (Boersma and Weenink, 2005), which includes spectrum and spectrogram analysis, pitch extraction and formant analysis, and an embedded scripting language for automation. It is available on Microsoft, Macintosh, and UNIX environments.

7.6 ADVANCED: ARTICULATORY AND GESTURAL PHONOLOGY

We saw in Sec. 7.3.1 that we could use distinctive features to capture generalizations across phone class. These generalizations were mainly articulatory (although some, like [strident] and the vowel height features, are primarily acoustic).

原书第 267 页

This idea that articulation underlies phonetic production is used in a more sophisticated way in articulatory phonology, in which the articulatory gesture is the underlying phonological abstraction (Browman and Goldstein, 1986, 1992). Articulatory gestures are defined as parameterized dynamical systems. Since speech production requires the coordinated actions of tongue, lips, glottis, etc, articulatory phonology represents a speech utterances as a sequence of potentially overlapping articulatory gestures. Fig. 7.29 shows the sequential sequence (or gestural score) required for the production of the word $ pawn $ [p aa n]. The lips first close, then the glottis opens, then the tongue body moves back toward the pharynx wall for the vowel [aa], the velum drops for the nasal sounds, and finally the tongue tip closes against the alveolar ridge. The lines in the diagram indicate gestures which are phased with respect to each other. With such a gestural representation, the nasality in the [aa] vowel is explained by the timing of the gestures; the velum drops before the tongue tip has quite closed.

Image
Figure 7.29 The gestural score for the word pawn as pronounced [p aɑːn], after Browman and Goldstein (1989) and Browman and Goldstein (1995).

The intuition behind articulatory phonology is that the gestural score is likely to be much better as a set of hidden states at capturing the continuous nature of speech than a discrete sequence of phones. In addition, using articulatory gestures as a basic unit can help in modeling the fine-grained effects of coarticulation of neighboring gestures that we will explore further when we introduce diphones (Sec. ??) and triphones (Sec. ??).

Computational implementations of articulatory phonology have recently appeared in speech recognition, using articulatory gestures rather than phones as the underlying representation or hidden variable. Since multiple articulators (tongue, lips, etc) can move simultaneously, using gestures as the hidden variable implies a multi-tier hidden representation. Fig. 7.30 shows the articulatory feature set used in the work of Livescu and Glass (2004) and Livescu (2005); Fig. 7.31 shows examples of how phones are mapped onto this feature set.

原书第 268 页
FeatureDescriptionvalue = meaning
LIP-LOCposition of lipsLAB = labial (neutral position); PRO = protruded (rounded); DEN = dental
LIP-OPENdegree of opening of lipsCL = closed; CR = critical (labial/labio-dental fricative); NA = narrow (e.g., [w], [uw]); WI = wide (all other sounds)
TT-LOClocation of tongue tipDEN = inter-dental ([th], [dh]); ALV = alveolar ([t], [n]); P-A = palato-alveolar ([sh]); RET = retroflex ([r])
TT-OPENdegree of opening of tongue tipCL = closed (stop); CR = critical (fricative); NA = narrow ([r], alveolar glide); M-N = medium-narrow;MID = medium;WI = wide
TB-LOClocation of tongue bodyPAL = palatal (e.g. [sh], [y]); VEL = velar (e.g., [k], [ng]); UVU = uvular (neutral position);PHA = pharyngeal (e.g. [aa])
TB-OPENdegree of opening of tongue bodyCL = closed (stop); CR = critical (e.g. fricated [g] in "legal"); NA = narrow (e.g. [y]); M-N = medium-narrow;MID = medium;WI = wide
VELstate of the velumCL = closed (non-nasal); OP = open (nasal)
GLOTstate of the glottisCL = closed (glottal stop); CR = critical (voiced); OP = open (voiceless)
Figure 7.30 Articulatory-phonology-based feature set from Livescu (2005).
phoneLIP-LOCLIP-OPENTT-LOCTT-OPENTB-LOCTB-OPENVELGLOT
aaLABWALVWPHAM-NCL(.9),OP(.1)CR
aeLABWALVWVELWCL(.9),OP(.1)CR
bLABCRALVMUVUWCLCR
fDENCRALVMVELMCLOP
nLABWALVCLUVUMOPCR
sLABWALVCRUVUMCLOP
uwPRONP-AWVELNCL(.9),OP(.1)CR
Figure 7.31 Livescu (2005): sample of mapping from phones to underlying target articulatory feature values. Note that some values are probabilistic.

7.7 SUMMARY

This chapter has introduced many of the important concepts of phonetics and computational phonetics.

  • We can represent the pronunciation of words in terms of units called phones. The standard system for representing phones is the International Phonetic Alphabet or IPA. The most common computational system for transcription of English is the ARPAbet, which conveniently uses ASCII symbols.
  • Phones can be described by how they are produced articulatorily by the vocal organs; consonants are defined in terms of their place and manner of articulation and voicing, vowels by their height, backness, and roundness.

• A phoneme is a generalization or abstraction over different phonetic realizations.

Allophonic rules express how a phoneme is realized in a given context.

  • Speech sounds can also be described acoustically. Sound waves can be described in terms of frequency, amplitude, or their perceptual correlates, pitch and loudness.
  • The spectrum of a sound describes its different frequency components. While some phonetic properties are recognizable from the waveform, both humans and machines rely on spectral analysis for phone detection.
原书第 269 页
  • A spectrogram is a plot of a spectrum over time. Vowels are described by characteristic harmonics called formants.
  • Pronunciation dictionaries are widely available, and used for both speech recognition and speech synthesis, including the CMU dictionary for English and CELEX dictionaries for English, German, and Dutch. Other dictionaries are available from the LDC.

• Phonetically transcribed corpora are a useful resource for building computational models of phone variation and reduction in natural speech.

BIBLIOGRAPHICAL AND HISTORICAL NOTES

The major insights of articulatory phonetics date to the linguists of 800–150 B.C. India. They invented the concepts of place and manner of articulation, worked out the glottal mechanism of voicing, and understood the concept of assimilation. European science did not catch up with the Indian phoneticians until over 2000 years later, in the late 19th century. The Greeks did have some rudimentary phonetic knowledge; by the time of Plato's Theaetetus and Cratylus, for example, they distinguished vowels from consonants, and stop consonants from continuants. The Stoics developed the idea of the syllable and were aware of phonotactic constraints on possible words. An unknown Icelandic scholar of the twelfth century exploited the concept of the phoneme, proposed a phonemic writing system for Icelandic, including diacritics for length and nasality. But his text remained unpublished until 1818 and even then was largely unknown outside Scandinavia (Robins, 1967). The modern era of phonetics is usually said to have begun with Sweet, who proposed what is essentially the phoneme in his Handbook of Phonetics (1877). He also devised an alphabet for transcription and distinguished between broad and narrow transcription, proposing many ideas that were eventually incorporated into the IPA. Sweet was considered the best practicing phonetician of his time; he made the first scientific recordings of languages for phonetic purposes, and advanced the state of the art of articulatory description. He was also infamously difficult to get along with, a trait that is well captured in Henry Higgins, the stage character that George Bernard Shaw modeled after him. The phoneme was first named by the Polish scholar Baudouin de Courtenay, who published his theories in 1894.

Students with further interest in transcription and articulatory phonetics should consult an introductory phonetics textbook such as Ladefoged (1993) or Clark and Yallop (1995). Pullum and Ladusaw (1996) is a comprehensive guide to each of the symbols and diacritics of the IPA. A good resource for details about reduction and other phonetic processes in spoken English is Shockey (2003). Wells (1982) is the definitive 3-volume source on dialects of English.

Many of the classic insights in acoustic phonetics had been developed by the late 1950's or early 1960's; just a few highlights include techniques like the sound spectrograph (Koenig et al., 1946), theoretical insights like the working out of the source-filter theory and other issues in the mapping between articulation and acoustics (Fant, 1970; Stevens et al., 1953; Stevens and House, 1955; Heinz and Stevens, 1961; Stevens and

原书第 270 页

House, 1961), the F1xF2 space of vowel formants Peterson and Barney (1952), the understanding of the phonetic nature of stress and the use of duration and intensity as cues (Fry, 1955), and a basic understanding of issues in phone perception Miller and Nicely (1955), Liberman et al. (1952). Lehiste (1967) is a collection of classic papers on acoustic phonetics.

Excellent textbooks on acoustic phonetics include Johnson (2003) and (Ladefoged, 1996). (Coleman, 2005) includes an introduction to computational processing of acoustics as well as other speech processing issues, from a linguistic perspective. (Stevens, 1998) lays out an influential theory of speech sound production. There are a wide variety of books that address speech from a signal processing and electrical engineering perspective. The ones with the greatest coverage of computational phonetics issues include (Huang et al., 2001), (O'Shaughnessy, 2000), and (Gold and Morgan, 1999). An excellent textbook on digital signal processing is Lyons (2004).

There are a number of software packages for acoustic phonetic analysis. Probably the most widely-used one is Praat (Boersma and Weenink, 2005).

Many phonetics papers of computational interest are to be found in the Journal of the Acoustical Society of America (JASA), Computer Speech and Language, and Speech Communication.

EXERCISES

7.1 Find the mistakes in the ARPAbet transcriptions of the following words:

a. “three” [dh r i] d. “study” [s t uh d i] g. “slight” [s l iy t]

b. “sing” [s i h n g] e. “though” [th ow]

c. “eyes” [ay s] f. “planning” [p pl aa n ih ng]

7.2 Translate the pronunciations of the following color words from the IPA into the ARPAbet (and make a note if you think you pronounce them differently than this!):

a. [rɛd] e. [blæk] i. [pjus]

b. [blu] f. [wait] j. [toup]

c. [grin] g. [ɔrnɪdʒ]

d. [jelou] h. [pɜpl]

7.3 Ira Gershwin's lyric for Let's Call the Whole Thing Off talks about two pronunciations of the word "either" (in addition to the tomato and potato example given at the beginning of the chapter. Transcribe Ira Gershwin's two pronunciations of "either" in the ARPAbet.

7.4 Transcribe the following words in the ARPAbet:

a. dark

b. suit

原书第 271 页

c. greasy

d. wash

e. water

7.5 Take a wavefile of your choice. Some examples are on the textbook website. Download the PRAAT software, and use it to transcribe the wavefiles at the word level, and into ARPAbet phones, using Praat to help you play pieces of each wavefile, and to look at the wavefile and the spectrogram.

7.6 Record yourself saying five of the English vowels: [aah], [eh], [ae], [iy], [uw]. Find F1 and F2 for each of your vowels.

原书第 272 页

Baayen, R. H., Piepenbrock, R., and Gulikers, L. (1995). The CELEX Lexical Database (Release 2) [CD-ROM]. Linguistic Data Consortium, University of Pennsylvania [Distributor], Philadelphia, PA.

Bell, A., Jurafsky, D., Fosler-Lussier, E., Girand, C., Gregory, M. L., and Gildea, D. (2003). Effects of disfluencies, predictability, and utterance position on word form variation in English conversation. Journal of the Acoustical Society of America, 113(2), 1001–1024.

Boersma, P. and Weenink, D. (2005). Praat: doing phonetics by computer (version 4.3.14). [Computer program]. Retrieved May 26, 2005, from http://www.praat.org/.

Bolinger, D. (1981). Two kinds of vowels, two kinds of rhythm. Indiana University Linguistics Club.

Browman, C. P. and Goldstein, L. (1986). Towards an articulatory phonology. Phonology Yearbook, 3, 219–252.

Browman, C. P. and Goldstein, L. (1992). Articulatory phonology: An overview. Phonética, 49, 155–180.

Browman, C. P. and Goldstein, L. (1989). Articulatory gestures as phonological units. Phonology, 6, 201–250.

Browman, C. P. and Goldstein, L. (1995). Dynamics and articulatory phonology. In Port, R. and v. Gelder, T. (Eds.), Mind as Motion: Explorations in the Dynamics of Cognition, pp. 175–193. MIT Press.

Bybee, J. L. (2000). The phonology of the lexicon: evidence from lexical diffusion. In Barlow, M. and Kemmer, S. (Eds.), Usage-based Models of Language, pp. 65–85. CSLI, Stanford.

Cedergren, H. J. and Sankoff, D. (1974). Variable rules: performance as a statistical reflection of competence. Language, 50(2), 333–355.

Chomsky, N. and Halle, M. (1968). The Sound Pattern of English. Harper and Row.

Clark, J. and Yallop, C. (1995). An Introduction to Phonetics and Phonology. Blackwell, Oxford. 2nd ed.

CMU (1993). The Carnegie Mellon Pronouncing Dictionary v0.1. Carnegie Mellon University.

Coleman, J. (2005). Introducing Speech and Language Processing. Cambridge University Press.

Fant, G. M. (1960). Acoustic Theory of Speech Production. Mouton.

Fox Tree, J. E. and Clark, H. H. (1997). Pronouncing “the” as “thee” to signal problems in speaking. Cognition, 62, 151–167.

Fry, D. B. (1955). Duration and intensity as physical correlates of linguistic stress. Journal of the Acoustical Society of America, 27, 765–768.

Godfrey, J., Holliman, E., and McDaniel, J. (1992). SWITCHBOARD: Telephone speech corpus for research and development. In IEEE ICASSP-92, San Francisco, pp. 517–520. IEEE.

Gold, B. and Morgan, N. (1999). Speech and Audio Signal Processing. Wiley Press.

Greenberg, S., Ellis, D., and Hollenback, J. (1996). Insights into spoken language gleaned from phonetic transcription of the Switchboard corpus. In ICSLP-96, Philadelphia, PA, pp. S24–27.

Gregory, M. L., Raymond, W. D., Bell, A., Fosler-Lussier, E., and Jurafsky, D. (1999). The effects of collocational strength and contextual predictability in lexical production. In CLS-99, pp. 151–166. University of Chicago, Chicago.

Guy, G. R. (1980). Variation in the group and the individual: The case of final stop deletion. In Labov, W. (Ed.), Locating Language in Time and Space, pp. 1–36. Academic.

Heinz, J. M. and Stevens, K. N. (1961). On the properties of voiceless fricative consonants. Journal of the Acoustical Society of America, 33, 589–596.

Huang, X., Acero, A., and Hon, H.-W. (2001). Spoken Language Processing: A Guide to Theory, Algorithm, and System Development. Prentice Hall, Upper Saddle River, NJ.

Johnson, K. (2003). Acoustic and Auditory Phonetics. Blackwell, Oxford. 2nd ed.

Keating, P. A., Byrd, D., Flemming, E., and Todaka, Y. (1994). Phonetic analysis of word and segment variation using the TIMIT corpus of American English. Speech Communication, 14, 131–142.

Kirchhoff, K., Fink, G. A., and Sagerer, G. (2002). Combining acoustic and articulatory feature information for robust speech recognition. Speech Communication, 37, 303–319.

Koenig, W., Dunn, H. K., Y., L., and Lacy (1946). The sound spectrograph. Journal of the Acoustical Society of America, 18, 19–49.

Labov, W. (1966). The Social Stratification of English in New York City. Center for Applied Linguistics, Washington, D.C.

Labov, W. (1972). The internal evolution of linguistic rules. In Stockwell, R. P. and Macaulay, R. K. S. (Eds.), Linguistic Change and Generative Theory, pp. 101–171. Indiana University Press, Bloomington.

Labov, W. (1975). The quantitative study of linguistic structure. Pennsylvania Working Papers on Linguistic Change and Variation v.1 no. 3. U.S. Regional Survey, Philadelphia, PA.

Labov, W. (1994). Principles of Linguistic Change: Internal Factors. Blackwell, Oxford.

Ladefoged, P. (1993). A Course in Phonetics. Harcourt Brace Jovanovich. Third Edition.

Ladefoged, P. (1996). Elements of Acoustic Phonetics. University of Chicago, Chicago, IL. Second Edition.

LDC (1995). COMLEX English Pronunciation Dictionary Version 0.2 (COMLEX 0.2). Linguistic Data Consortium.

Lehiste, I. (Ed.). (1967). Readings in Acoustic Phonetics. MIT Press.

Li, A., Zheng, F., Byrne, W., Fung, P., Kamm, T., Yi, L., Song, Z., Ruh, U., Venkataramani, V., and Chen, X. (2000). CASS: A phonetically transcribed corpus of mandarin spontaneous speech. In ICSLP-00, Beijing, China, pp. 485–488.

原书第 273 页

Liberman, A. M., Delattre, P. C., and Cooper, F. S. (1952). The role of selected stimulus variables in the perception of the unvoiced stop consonants. American Journal of Psychology, 65, 497–516.

Livescu, K. (2005). Feature-Based Pronunciation Modeling for Automatic Speech Recognition. Ph.D. thesis, Massachusetts Institute of Technology.

Livescu, K. and Glass, J. (2004). Feature-based pronunciation modeling with trainable asynchrony probabilities. In ICSLP-04, Jeju, South Korea.

Lyons, R. G. (2004). Understanding Digital Signal Processing. Prentice Hall, Upper Saddle River, NJ. 2nd ed.

Miller, C. A. (1998). Pronunciation modeling in speech synthesis. Tech. rep. IRCS 98–09, University of Pennsylvania Institute for Research in Cognitive Science, Philadelphia, PA.

Miller, G. A. and Nicely, P. E. (1955). An analysis of perceptual confusions among some English consonants. Journal of the Acoustical Society of America, 27, 338–352.

Morgan, N. and Fosler-Lussier, E. (1989). Combining multiple estimators of speaking rate. In IEEE ICASSP-89.

Neu, H. (1980). Ranking of constraints on /t,d/ deletion in American English: A statistical analysis. In Labov, W. (Ed.), Locating Language in Time and Space, pp. 37–54. Academic Press.

NIST (1990). TIMIT Acoustic-Phonetic Continuous Speech Corpus. National Institute of Standards and Technology Speech Disc 1-1.1. NIST Order No. PB91-505065.

O'Shaughnessy, D. (2000). Speech Communications: Human and Machine. IEEE Press, New York. 2nd ed.

Peterson, G. E. and Barney, H. L. (1952). Control methods used in a study of the vowels. Journal of the Acoustical Society of America, 24, 175–184.

Pullum, G. K. and Ladusaw, W. A. (1996). Phonetic Symbol Guide. University of Chicago, Chicago, IL. Second Edition.

Rand, D. and Sankoff, D. (1990). Goldvarb: A variable rule application for the macintosh. Available at http://www.crm.umontreal.ca/sankoff/GoldVarb_Eng.html.

Rhodes, R. A. (1992). Flapping in American English. In Dressler, W. U., Prinzhorn, M., and Rennison, J. (Eds.), Proceedings of the 7th International Phonology Meeting, pp. 217–232. Rosenberg and Sellier, Turin.

Robins, R. H. (1967). A Short History of Linguistics. Indiana University Press, Bloomington.

Seneff, S. and Zue, V. W. (1988). Transcription and alignment of the TIMIT database. In Proceedings of the Second Symposium on Advanced Man-Machine Interface through Spoken Language, Oahu, Hawaii.

Shockey, L. (2003). Sound Patterns of Spoken English. Blackwell, Oxford.

Shoup, J. E. (1980). Phonological aspects of speech recognition. In Lea, W. A. (Ed.), Trends in Speech Recognition, pp. 125–138. Prentice-Hall.

Stevens, K. N., Kasowski, S., and Fant, G. M. (1953). An electrical analog of the vocal tract. Journal of the Acoustical Society of America, 25(4), 734–742.

Stevens, K. N. (1998). Acoustic Phonetics. MIT Press.

Stevens, K. N. and House, A. S. (1955). Development of a quantitative description of vowel articulation. Journal of the Acoustical Society of America, 27, 484–493.

Stevens, K. N. and House, A. S. (1961). An acoustical theory of vowel production and some of its implications. Journal of Speech and Hearing Research, 4, 303–320.

Stevens, S. S. and Volkmann, J. (1940). The relation of pitch to frequency: A revised scale. The American Journal of Psychology, 53(3), 329–353.

Stevens, S. S., Volkmann, J., and Newman, E. B. (1937). A scale for the measurement of the psychological magnitude pitch. Journal of the Acoustical Society of America, 8, 185–190.

Sweet, H. (1877). A Handbook of Phonetics. Clarendon Press, Oxford.

Wald, B. and Shopen, T. (1981). A researcher's guide to the sociolinguistic variable (ING). In Shopen, T. and Williams, J. M. (Eds.), Style and Variables in English, pp. 219–249. Winthrop Publishers.

Wells, J. C. (1982). Accents of English. Cambridge University Press.

Wolfram, W. A. (1969). A Sociolinguistic Description of Detroit Negro Speech. Center for Applied Linguistics, Washington, D.C.

Zappa, F. and Zappa, M. U. (1982). Valley girl. From Frank Zappa album Ship Arriving Too Late To Save A Drowning Witch.

Zwicky, A. (1972). On Casual Speech. In CLS-72, pp. 607–615. University of Chicago.

原书第 274 页

8

SPEECH SYNTHESIS

And computers are getting smarter all the time: Scientists tell us that soon they will be able to talk to us. (By ‘they’ I mean ‘computers’: I doubt scientists will ever be able to talk to us.)

Dave Barry

In Vienna in 1769, Wolfgang von Kempelen built for the Empress Maria Theresa the famous Mechanical Turk, a chess-playing automaton consisting of a wooden box filled with gears, and a robot mannequin sitting behind the box who played chess by moving pieces with his mechanical arm. The Turk toured Europe and the Americas for decades, defeating Napoleon Bonaparte and even playing Charles Babbage. The Mechanical Turk might have been one of the early successes of artificial intelligence if it were not for the fact that it was, alas, a hoax, powered by a human chesplayer hidden inside the box.

What is perhaps less well-known is that von Kempelen, an extraordinarily prolific inventor, also built between 1769 and 1790 what is definitely not a hoax: the first full-sentence speech synthesizer. His device consisted of a bellows to simulate the lungs, a rubber mouthpiece and a nose aperture, a reed to simulate the vocal folds, various whistles for each of the fricatives. and a small auxiliary bellows to provide the puff of air for plosives. By moving levers with both hands, opening and closing various openings, and adjusting the flexible leather ‘vocal tract’, different consonants and vowels could be produced.

More than two centuries later, we no longer build our speech synthesizers out of wood, leather, and rubber, nor do we need trained human operators. The modern task of speech synthesis, also called text-to-speech or TTS, is to produce speech (acoustic waveforms) from text input.

Modern speech synthesis has a wide variety of applications. Synthesizers are used, together with speech recognizers, in telephone-based conversational agents that conduct dialogues with people (see Ch. 23). Synthesizer are also important in non-conversational applications that speak to people, such as in devices that read out loud for the blind, or in video games or children's toys. Finally, speech synthesis can be used to speak for sufferers of neurological disorders, such as astrophysicist Steven Hawking who, having lost the use of his voice due to ALS, speaks by typing to a speech synthe-

原书第 275 页

sizer and having the synthesizer speak out the words. State of the art systems in speech synthesis can achieve remarkably natural speech for a very wide variety of input situations, although even the best systems still tend to sound wooden and are limited in the voices they use.

The task of speech synthesis is to map a text like the following:

(8.1) PG&E will file schedules on April 20.

to a waveform like the following:

Image

Speech synthesis systems perform this mapping in two steps, first converting the input text into a phonemic internal representation and then converting this internal representation into a waveform. We will call the first step text analysis and the second step waveform synthesis (although other names are also used for these steps).

A sample of the internal representation for this sentence is shown in Fig. 8.1. Note that the acronym PG&E is expanded into the words P G AND E, the number 20 is expanded into twentieth, a phone sequence is given for each of the words, and there is also prosodic and phrasing information (the *'s) which we will define later.

PGANDEWILLFILESCHEDULESONAPRILTWENTIETH
p[iyjh]iyaendiywihlfaylskehjhaxlzaaneyprihltwehntiyaxth
Figure 8.1 Intermediate output for a unit selection synthesizer for the sentence PG&E will file schedules on April 20. The numbers and acronyms have been expanded, words have been converted into phones, and prosodic features have been assigned.

While text analysis algorithms are relatively standard, there are three widely different paradigms for waveform synthesis: concatenative synthesis, formant synthesis, and articulatory synthesis. The architecture of most modern commercial TTS systems is based on concatenative synthesis, in which samples of speech are chopped up, stored in a database, and combined and reconfigured to create new sentences. Thus we will focus on concatenative synthesis for most of this chapter, although we will briefly introduce formant and articulatory synthesis at the end of the chapter.

HOURGLASS METAPHOR

Fig. 8.2 shows the TTS architecture for concatenative unit selection synthesis, using the two-step hourglass metaphor of Taylor (2008). In the following sections, we'll examine each of the components in this architecture.

8.1 TEXT NORMALIZATION

In order to generate a phonemic internal representation, raw text first needs to be preprocessed or normalized in a variety of ways. We'll need to break the input text into

原书第 276 页
Image
Figure 8.2 Architecture for the unit selection (concatenative) architecture for speech synthesis.

sentences, and deal with the idiosyncracies of abbreviations, numbers, and so on. Consider the difficulties in the following text drawn from the Enron corpus (Klimt and Yang, 2004):

He said the increase in credit limits helped B.C. Hydro achieve record net income of about $1 billion during the year ending March 31. This figure does not include any write-downs that may occur if Powerex determines that any of its customer accounts are not collectible. Cousins, however, was insistent that all debts will be collected: “We continue to pursue monies owing and we expect to be paid for electricity we have sold.”

SENTENCE TOKENIZATION

The first task in text normalization is sentence tokenization. In order to segment this paragraph into separate utterances for synthesis, we need to know that the first sentence ends at the period after March 31, not at the period of B.C. We also need to know that there is a sentence ending at the word collected, despite the punctuation being a colon rather than a period. The second normalization task is dealing with non-standard words. Non-standard words include number, acronyms, abbreviations, and so on. For example, March 31 needs to be pronounced March thirty-first, not March three one; $1 billion needs to be pronounced one billion dollars, with the word dollars appearing after the word billion.

原书第 277 页
← 7.4.5 Spectra and the Frequency Domain8.1.1 Sentence Tokenization →