← 学习库 Speech and Language Processing 本册目录

(24.1) Turn-taking Rule. At each TRP of each turn:

a. If during this turn the current speaker has selected A as the next speaker then A must speak next.

b. If the current speaker does not select the next speaker, any other speaker may take the next turn.

c. If no one else takes the next turn, the current speaker may take the next turn.

There are a number of important implications of rule (24.1) for dialogue modeling. First, subrule (24.1a) implies that there are some utterances by which the speaker specifically selects who the next speaker will be. The most obvious of these are questions, in which the speaker selects another speaker to answer the question. Two-part

原书第 939 页

structures like QUESTION-ANSWER are called adjacency pairs (Schegloff, 1968) or dialogic pair (Harris, 2005). Other adjacency pairs include GREETING followed by GREETING, COMPLIMENT followed by DOWNPLAYER, REQUEST followed by GRANT. We will see that these pairs and the dialogue expectations they set up will play an important role in dialogue modeling.

(24.2) A: Is there something bothering you or not?

Subrule (24.1a) also has an implication for the interpretation of silence. While silence can occur after any turn, silence in between the two parts of an adjacency pair is significant silence. For example Levinson (1983) notes this example from Atkinson and Drew (1979); pause lengths are marked in parentheses (in seconds):

(1.0)

(1.5)

A: Yes or no?

PERFORMATIVE

A: Eh?

B: No.

Since A has just asked B a question, the silence is interpreted as a refusal to respond, or perhaps a dispreferred response (a response, like saying “no” to a request, which is stigmatized). By contrast, silence in other places, for example a lapse after a speaker finishes a turn, is not generally interpretable in this way. These facts are relevant for user interface design in spoken dialogue systems; users are disturbed by the pauses in dialogue systems caused by slow speech recognizers (Yankelovich et al., 1995).

Another implication of (24.1) is that transitions between speakers don't occur just anywhere; the transition-relevance places where they tend to occur are generally at utterance boundaries. Recall from Ch. 12 that spoken utterances differ from written sentences in a number of ways. They tend to be shorter, are more likely to be single clauses or even just single words, the subjects are usually pronouns rather than full lexical noun phrases, and they include filled pauses and repairs. A hearer must take all this (and other cues like prosody) into account to know where to begin talking.

← 24.1.1 Turns and Turn-Taking24.1.2 Language as Action: Speech Acts →