← 学习库 Speech and Language Processing 本册目录

7.2.4 Vowels

Like consonants, vowels can be characterized by the position of the articulators as they are made. The three most relevant parameters for vowels are what is called vowel height, which correlates roughly with the height of the highest part of the tongue, vowel frontness or backness, which indicates whether this high point is toward the front or back of the oral tract, and the shape of the lips (rounded or not). Fig. 7.5 shows the position of the tongue for different vowels.

Image
Figure 7.5 Positions of the tongue for three English vowels, high front [iy], low front [ae] and high back [uw]; tongue positions modeled after Ladefoged (1996).

In the vowel [iy], for example, the highest point of the tongue is toward the front of the mouth. In the vowel [uw], by contrast, the high-point of the tongue is located toward the back of the mouth. Vowels in which the tongue is raised toward the front are called front vowels; those in which the tongue is raised toward the back are called

原书第 245 页

back vowels. Note that while both [ih] and [eh] are front vowels, the tongue is higher for [ih] than for [eh]. Vowels in which the highest point of the tongue is comparatively high are called high vowels; vowels with mid or low values of maximum tongue height are called mid vowels or low vowels, respectively.

Image
Figure 7.6 Qualities of English vowels (after Ladefoged (1993)).

Fig. 7.6 shows a schematic characterization of the vowel height of different vowels. It is schematic because the abstract property height only correlates roughly with actual tongue positions; it is in fact a more accurate reflection of acoustic facts. Note that the chart has two kinds of vowels: those in which tongue height is represented as a point and those in which it is represented as a vector. A vowel in which the tongue position changes markedly during the production of the vowel is a diphthong. English is particularly rich in diphthongs.

SYLLABLE

The second important articulatory dimension for vowels is the shape of the lips. Certain vowels are pronounced with the lips rounded (the same lip shape used for whistling). These rounded vowels include [uw], [ao], and [ow].

Syllables

Consonants and vowels combine to make a syllable. There is no completely agreed-upon definition of a syllable; roughly speaking a syllable is a vowel-like (or sonorant) sound together with some of the surrounding consonants that are most closely associated with it. The word dog has one syllable, [d aa g], while the word catnip has two syllables, [k ae t] and [n ih p]. We call the vowel at the core of a syllable the nucleus. The optional initial consonant or set of consonants is called the onset. If the onset has more than one consonant (as in the word strike [s t r ay k]), we say it has a complex onset. The coda. is the optional consonant or sequence of consonants following the nucleus. Thus [d] is the onset of dog, while [g] is the coda. The rime. or rhyme. is the nucleus plus coda. Fig. 7.7 shows some sample syllable structures.

原书第 246 页
Image
Figure 7.7 Syllable structure of ham, green, eggs. $ \sigma $=syllable.

The task of automatically breaking up a word into syllables is called syllabification, and will be discussed in Sec. ??.

Syllable structure is also closely related to the phonotactics of a language. The term phonotactics means the constraints on which phones can follow each other in a language. For example, English has strong constraints on what kinds of consonants can appear together in an onset; the sequence [zdr], for example, cannot be a legal English syllable onset. Phonotactics can be represented by listing constraints on fillers of syllable positions, or by creating a finite-state model of possible phone sequences. It is also possible to create a probabilistic phonotactics, by training N-gram grammars on phone sequences.

Lexical Stress and Schwa

In a natural sentence of American English, certain syllables are more prominent than others. These are called accented syllables, and the linguistic marker associated with this prominence is called a pitch accent. Words or syllables which are prominent are said to bear (be associated with) a pitch accent. Pitch accent is also sometimes referred to as sentence stress, although sentence stress can instead refer to only the most prominent accent in a sentence.

Accented syllables may be prominent by being louder, longer, by being associated with a pitch movement, or by any combination of the above. Since accent plays important roles in meaning, understanding exactly why a speaker chooses to accent a particular syllable is very complex, and we will return to this in detail in Sec. ??. But one important factor in accent is often represented in pronunciation dictionaries. This factor is called lexical stress. The syllable that has lexical stress is the one that will be louder or longer if the word is accented. For example the word parsley is stressed in its first syllable, not its second. Thus if the word parsley receives a pitch accent in a sentence, it is the first syllable that will be stronger.

In IPA we write the symbol ['] before a syllable to indicate that it has lexical stress (e.g. ['par.sli]). This difference in lexical stress can affect the meaning of a word. For example the word content can be a noun or an adjective. When pronounced in isolation the two senses are pronounced differently since they have different stressed syllables (the noun is pronounced ['kən.tent] and the adjective [kən.tent]).

Vowels which are unstressed can be weakened even further to reduced vowels. The most common reduced vowel is schwa ([ax]). Reduced vowels in English don't have their full form; the articulatory gesture isn't as complete as for a full vowel. As a result

原书第 247 页

the shape of the mouth is somewhat neutral; the tongue is neither particularly high nor particularly low. For example the second vowel in parakeet is a schwa: [p a e r a x k i y t].

While schwa is the most common reduced vowel, it is not the only one, at least not in some dialects. Bolinger (1981) proposed that American English had three reduced vowels: a reduced mid vowel [ə], a reduced front vowel [i], and a reduced rounded vowel [ə]. The full ARPAbet includes two of these, the schwa [ax] and [ix] ([i]), as well as [axr] which is an r-colored schwa (often called schwar), although [ix] is generally dropped in computational applications (Miller, 1998), and [ax] and [ix] are falling together in many dialects of English Wells (1982, p. 167–168).

Not all unstressed vowels are reduced; any vowel, and diphthongs in particular can retain their full quality even in unstressed position. For example the vowel [iy] can appear in stressed position as in the word eat [iy t] or in unstressed position in the word carry [k ae riy].

Some computational ARPAbet lexicons mark reduced vowels like schwa explicitly. But in general predicting reduction requires knowledge of things outside the lexicon (the prosodic context, rate of speech, etc, as we will see the next section). Thus other ARPAbet versions mark stress but don't mark how stress affects reduction. The CMU dictionary (CMU, 1993), for example, marks each vowel with the number 0 (unstressed) 1 (stressed), or 2 (secondary stress). Thus the word counter is listed as [K AW1 N T ER0], and the word table as [T EY1 B AH0 L]. Secondary stress is defined as a level of stress lower than primary stress, but higher than an unstressed vowel, as in the word dictionary [D IH1 K SH AH0 N EH2 R IY0].

We have mentioned a number of potential levels of prominence: accented, stressed, secondary stress, full vowel, and reduced vowel. It is still an open research question exactly how many levels are appropriate. Very few computational systems make use of all five of these levels, most using between one and three. We return to this discussion when we introduce prosody in more detail in Sec. ??.

7.3 PHONOLOGICAL CATEGORIES AND PRONUNCIATION VARIATION

'Scuse me, while I kiss the sky

Jimi Hendrix, Purple Haze

'Scuse me, while I kiss this guy

Common mis-hearing of same lyrics

If each word was pronounced with a fixed string of phones, each of which was pronounced the same in all contexts and by all speakers, the speech recognition and speech synthesis tasks would be really easy. Alas, the realization of words and phones varies massively depending on many factors. Fig. 7.8 shows a sample of the wide variation in pronunciation in the words because and about from the hand-transcribed Switchboard corpus of American English telephone conversations (Greenberg et al., 1996).

How can we model and predict this extensive variation? One useful tool is the assumption that what is mentally represented in the speaker's mind are abstract cate-

原书第 248 页
becauseabout
ARPAbet%ARPAbet%ARPAbet%ARPAbet%
b iy k ah z27%k s2%ax b aw32%b ae3%
b ix k ah z14%k ix z2%ax b aw t16%b aw t3%
k ah z7%k ih z2%b aw9%ax b aw dx3%
k ax z5%b iy k ah zh2%ix b aw8%ax b ae3%
b ix k ax z4%b iy k ah s2%ix b aw t5%b aa3%
b ih k ah z3%b iy k ah2%ix b ae4%b ae dx3%
b ax k ah z3%b iy k aa z2%ax b ae dx3%ix b aw dx2%
k uh z2%ax z2%b aw dx3%ix b aa t2%
Figure 7.8 The 16 most common pronunciations of because and about from the hand-transcribed Switchboard corpus of American English conversational telephone speech (Godfrey et al., 1992; Greenberg et al., 1996).

gories rather than phones in all their gory phonetic detail. For example consider the different pronunciations of [t] in the words tunafish and starfish. The [t] of tunafish is aspirated. Aspiration is a period of voicelessness after a stop closure and before the onset of voicing of the following vowel. Since the vocal cords are not vibrating, aspiration sounds like a puff of air after the [t] and before the vowel. By contrast, a [t] following an initial [s] is unaspirated; thus the [t] in starfish ([s t aa r f ih sh]) has no period of voicelessness after the [t] closure. This variation in the realization of [t] is predictable: whenever a [t] begins a word or unreduced syllable in English, it is aspirated. The same variation occurs for [k]; the [k] of sky is often mis-heard as [g] in Jimi Hendrix's lyrics because [k] and [g] are both unaspirated. $ ^{2} $

There are other contextual variants of [t]. For example, when [t] occurs between two vowels, particularly when the first is stressed, it is often pronounced as a tap. Recall that a tap is a voiced sound in which the top of the tongue is curled up and back and struck quickly against the alveolar ridge. Thus the word buttercup is usually pronounced [b ah dx axr k uh p] rather than [b ah t axr k uh p]. Another variant of [t] occurs before the dental consonant [th]. Here the [t] becomes dentalized (IPA [t]). That is, instead of the tongue forming a closure against the alveolar ridge, the tongue touches the back of the teeth.

In both linguistics and in speech processing, we use abstract classes to capture the similarity among all these [t]s. The simplest abstract class is called the phoneme, and its different surface realizations in different contexts are called allophones. We traditionally write phonemes inside slashes. So in the above examples, /t/ is a phoneme whose allophones include (in IPA) [tʰ], [r], and [t]. Fig. 7.9 summarizes a number of allophones of /t/. In speech synthesis and recognition, we use phonesets like the ARPAbet to approximate this idea of abstract phoneme units, and represent pronunciation lexicons using ARPAbet phones. For this reason, the allophones listed in Fig. 7.1 tend to be used for narrow transcriptions for analysis purposes, and less often used in speech recognition or synthesis systems.

原书第 249 页

| IPA | ARPABet | Description | Environment | Example |

| --- | --- | --- | --- | --- |

| $ t^{h} $ | [t] | aspirated | in initial position | toucan |

| t | | unaspirated | after [s] or in reduced syllables | starfish |

| ? | [q] | glottal stop | word-finally or after vowel before [n] | kitten |

| ?t | [qt] | glottal stop t | sometimes word-finally | cat |

| r | [dx] | tap | between vowels | butter |

| $ t^{r} $ | [tcl] | unreleased t | before consonants or word-finally | fruitcake |

| t | | dental t | before dental consonants ([ \theta ]) | eighth |

| | | deleted t | sometimes word-finally | past |

Figure 7.9 Some allophones of /t/ in General American English.

Variation is even more common than Fig. 7.9 suggests. One factor influencing variation is that the more natural and colloquial speech becomes, and the faster the speaker talks, the more the sounds are shortened, reduced and generally run together. This phenomena is known as reduction or hypoarticulation. For example assimilation is the change in a segment to make it more like a neighboring segment. The dentalization of [t] to ([t]) before the dental consonant [θ] is an example of assimilation. A common type of assimilation cross-linguistically is palatalization, when the constriction for a segment moves closer to the palate than it normally would, because the following segment is palatal or alveolo-palatal. In the most common cases, /s/ becomes [sh], /z/ becomes [zh], /t/ becomes [ch] and /d/ becomes [jh], We saw one case of palatalization in Fig. 7.8 in the pronunciation of because as [b i y k a h zh], because the following word was you've. The lemma you (you, your, you've, and you'd) is extremely likely to cause palatalization in the Switchboard corpus.

Deletion is quite common in English speech. We saw examples of deletion of final t/ above, in the words about and it. Deletion of final t/ and /d/ has been extensively studied. /d/ is more likely to be deleted than /t/, and both are more likely to be deleted before a consonant (Labov, 1972). Fig. 7.10 shows examples of palatalization and final t/d deletion from the Switchboard corpus.

PalatalizationFinal t/d Deletion
PhraseLexicalReducedPhraseLexicalReduced
set yours eh t y ow rs eh ch erfind himf ay n d h ih mf ay n ix m
not yetn aa t y eh tn aa ch eh tand weae n d w iyeh n w iy
did youd ih d y uwd ih jh y ahdraft thed r ae f t dh iyd r ae f dh iy
Figure 7.10 Examples of palatalization and final t/d/ deletion from the Switchboard corpus. Some of the t/d examples may have glottalization instead of being completely deleted.
← 7.2.3 Consonants: Manner of Articulation7.3.1 Phonetic Features →