← 学习库 Speech and Language Processing 本册目录

24.7.2 The Intentional Structure of Dialogue

In Sec. ?? we introduced the idea that the segments of a discourse are related by coherence relations like Explanation or Elaboration which describe the informational relation between discourse segments. The BDI approach to utterance interpretation gives

原书第 979 页

rise to another view of coherence which is particularly relevant for dialogue, the intentional approach (Grosz and Sidner, 1986). According to this approach, what makes a dialogue coherent is its intentional structure, the plan-based intentions of the speaker underlying each utterance.

These intentions are instantiated in the model by assuming that each discourse has an underlying purpose held by the person who initiates it, called the discourse purpose (DP). Each discourse segment within the discourse has a corresponding purpose, a discourse segment purpose (DSP), which has a role in achieving the overall DP. Possible DPs/DSPs include intending that some agent intend to perform some physical task or that some agent believe some fact.

As opposed to the larger sets of coherence relations used in informational accounts of coherence, Grosz and Sidner propose only two such relations: dominance and satisfaction-precedence. DSP $ _{1} $ dominates DSP $ _{2} $ if satisfying DSP $ _{2} $ is intended to provide part of the satisfaction of DSP $ _{1} $. DSP $ _{1} $ satisfaction-precedes DSP $ _{2} $ if DSP $ _{1} $ must be satisfied before DSP $ _{2} $.

| $ C_{1} $: | I need to travel in May. |

| --- | --- |

| $ A_{1} $: | And, what day in May did you want to travel? |

| $ C_{2} $: | OK uh I need to be there for a meeting that's from the 12th to the 15th. |

| $ A_{2} $: | And you're flying into what city? |

| $ C_{3} $: | Seattle. |

| $ A_{3} $: | And what time would you like to leave Pittsburgh? |

| $ C_{4} $: | Uh hmm I don't think there's many options for non-stop. |

| $ A_{4} $: | Right. There's three non-stops today. |

| $ C_{5} $: | What are they? |

| $ A_{5} $: | The first one departs PGH at 10:00am arrives Seattle at 12:05 their time. The second flight departs PGH at 5:55pm, arrives Seattle at 8pm. And the last flight departs PGH at 8:15pm arrives Seattle at 10:28pm. |

| $ C_{6} $: | OK I'll take the 5ish flight on the night before on the 11th. |

| $ A_{6} $: | On the 11th? OK. Departing at 5:55pm arrives Seattle at 8pm, U.S. Air flight 115. |

| $ C_{7} $: | OK. |

Figure 24.23 A fragment from a telephone conversation between a client (C) and travel agent (A) (repeated from Fig. 24.4).

Consider the dialogue between a client (C) and a travel agent (A) that we saw earlier, repeated here in Fig. 24.23. Collaboratively, the caller and agent successfully identify a flight that suits the caller's needs. Achieving this joint goal requires that a top-level discourse intention be satisfied, listed as I1 below, in addition to several intermediate intentions that contributed to the satisfaction of I1, listed as I2-I5:

I1: (Intend C (Intend A (A find a flight for C)))

I2: (Intend A (Intend C (Tell C A departure date)))

I3: (Intend A (Intend C (Tell C A destination city)))

I4: (Intend A (Intend C (Tell C A departure time)))

原书第 980 页

I5: (Intend C (Intend A (A find a nonstop flight for C)))

Intentions I2–I5 are all subordinate to intention I1, as they were all adopted to meet preconditions for achieving intention I1. This is reflected in the dominance relationships below:

I1 dominates I2 ∧ I1 dominates I3 ∧ I1 dominates I4 ∧ I1 dominates I5

Furthermore, intentions I2 and I3 needed to be satisfied before intention I5, since the agent needed to know the departure date and destination in order to start listing nonstop flights. This is reflected in the satisfaction-precedence relationships below:

I2 satisfaction-precedes I5 $ \land $ I3 satisfaction-precedes I5

The dominance relations give rise to the discourse structure depicted in Figure 24.24. Each discourse segment is numbered in correspondence with the intention number that serves as its DP/DSP.

| DS1 | | | |

| --- | --- | --- | --- |

| $ C_{1} $ | DS $ _{2} $ | DS $ _{3} $ | DS $ _{4} $ |

| DS $ _{5} $ | | | |

| | | | |

Figure 24.24 Discourse Structure of the Flight Reservation Dialogue

Intentions and their relationships give rise to a coherent discourse based on their role in the overall plan that the caller is inferred to have. We assume that the caller and agent have the plan BOOK-FLIGHT described on page 43. This plan requires that the agent know the departure time and date and so on. As we discussed above, the agent can use the REQUEST-INFO action scheme from page 44 to ask the user for this information.

Subsidiary discourse segments are also called subdialogues; DS2 and DS3 in particular are information-sharing (Chu-Carroll and Carberry, 1998) knowledge precondition subdialogues (Lochbaum et al., 1990; Lochbaum, 1998), since they are initiated by the agent to help satisfy preconditions of a higher-level goal.

Algorithms for inferring intentional structure in dialogue work similarly to algorithms for inferring dialogue acts, either employing the BDI model (e.g., Litman, 1985; Grosz and Sidner, 1986; Litman and Allen, 1987; Carberry, 1990; Passonneau and Litman, 1993; Chu-Carroll and Carberry, 1998), or machine learning architectures based on cue phrases (Reichman, 1985; Grosz and Sidner, 1986; Hirschberg and Litman, 1993), prosody (Hirschberg and Pierrehumbert, 1986; Grosz and Hirschberg, 1992; Pierrehumbert and Hirschberg, 1990; Hirschberg and Nakatani, 1996), and other cues.

24.8 SUMMARY

Conversational agents are a crucial speech and language processing application that are already widely used commercially. Research on these agents relies crucially on an

原书第 981 页

understanding of human dialogue or conversational practices.

  • Dialogue systems generally have 5 components: speech recognition, natural language understanding, dialogue management, natural language generation, and speech synthesis. They may also have a task manager specific to the task domain.
  • Dialogue architectures for conversational agents include finite-state systems, frame-based production systems, and advanced systems such as information-state, Markov Decision Processes, and BDI (belief-desire-intention) models.

Turn-taking, grounding, conversational structure, implicature, and initiative are crucial human dialogue phenomena that must also be dealt with in conversational agents.

  • Speaking in dialogue is a kind of action; these acts are referred to as speech acts or dialogue acts. Models exist for generating and interpreting these acts.

BIBLIOGRAPHICAL AND HISTORICAL NOTES

Early work on speech and language processing had very little emphasis on the study of dialogue. The dialogue manager for the simulation of the paranoid agent PARRY (Colby et al., 1971), was a little more complex. Like ELIZA, it was based on a production system, but where ELIZA's rules were based only on the words in the user's previous sentence, PARRY's rules also rely on global variables indicating its emotional state. Furthermore, PARRY's output sometimes makes use of script-like sequences of statements when the conversation turns to its delusions. For example, if PARRY's anger variable is high, he will choose from a set of "hostile" outputs. If the input mentions his delusion topic, he will increase the value of his fear variable and then begin to express the sequence of statements related to his delusion.

The appearance of more sophisticated dialogue managers awaited the better understanding of human-human dialogue. Studies of the properties of human-human dialogue began to accumulate in the 1970's and 1980's. The Conversation Analysis community (Sacks et al., 1974; Jefferson, 1984; Schegloff, 1982) began to study the interactional properties of conversation. Grosz's (1977) dissertation significantly influenced the computational study of dialogue with its introduction of the study of dialogue structure, with its finding that "task-oriented dialogues have a structure that closely parallels the structure of the task being performed" (p. 27), which led to her work on intentional and attentional structure with Sidner. Lochbaum et al. (2000) is a good recent summary of the role of intentional structure in dialogue. The BDI model integrating earlier AI planning work (Fikes and Nilsson, 1971) with speech act theory (Austin, 1962; Gordon and Lakoff, 1971; Searle, 1975a) was first worked out by Cohen and Perrault (1979), showing how speech acts could be generated, and Perrault and Allen (1980) and Allen and Perrault (1980), applying the approach to speech-act interpretation. Simultaneous work on a plan-based model of understanding was developed by Wilensky (1983) in the Schankian tradition.

原书第 982 页

Probabilistic models of dialogue act interpretation were informed by linguistic work which focused on the discourse meaning of prosody (Sag and Liberman, 1975; Pierrehumbert, 1980), by Conversation Analysis work on microgrammar (e.g. Goodwin, 1996), by work such as Hinkelman and Allen (1989), who showed how lexical and phrasal cues could be integrated into the BDI model, and then worked out at a number of speech and dialogue labs in the 1990's (Waibel, 1988; Daly and Zue, 1992; Kompe et al., 1993; Nagata and Morimoto, 1994; Woszczyna and Waibel, 1994; Reithinger et al., 1996; Kita et al., 1996; Warnke et al., 1997; Chu-Carroll, 1998; Stolcke et al., 1998; Taylor et al., 1998; Stolcke et al., 2000).

Modern dialogue systems drew on research at many different labs in the 1980's and 1990's. Models of dialogue as collaborative behavior were introduced in the late 1980's and 1990's, including the ideas of common ground (Clark and Marshall, 1981), reference as a collaborative process (Clark and Wilkes-Gibbs, 1986), and models of joint intentions (Levesque et al., 1990), and shared plans (Grosz and Sidner, 1980). Related to this area is the study of initiative in dialogue, studying how the dialogue control shifts between participants (Walker and Whittaker, 1990; Smith and Gordon, 1997; Chu-Carroll and Brown, 1997).

A wide body of dialogue research came out of AT&T and Bell Laboratories around the turn of the century, including much of the early work on MDP dialogue systems as well as fundamental work on cue-phrases, prosody, and rejection and confirmation. Work on dialogue acts and dialogue moves drew from a number of sources, including HCRC's Map Task (Carletta et al., 1997b), and the work of James Allen and his colleagues and students, for example Hinkelman and Allen (1989), showing how lexical and phrasal cues could be integrated into the BDI model of speech acts, and Traum (2000), Traum and Hinkelman (1992), and from Sadek (1991).

Much recent academic work in dialogue focuses on multimodal applications (Johnston et al., 2007; Niekrasz and Purver, 2006, inter alia), on the information-state model (Traum and Larsson, 2003, 2000) or on reinforcement learning architectures including POMDPs (Roy et al., 2000; Young, 2002; Lemon et al., 2006; Williams and Young, 2005, 2000). Work in progress on MDPs and POMDPs focuses on computational complexity (they currently can only be run on quite small domains with limited numbers of slots), and on improving simulations to make them more reflective of true user behavior. Alternative algorithms include SMDPs (Cuayáhuitl et al., 2007). See Russell and Norvig (2002) and Sutton and Barto (1998) for a general introduction to reinforcement learning.

Recent years have seen the widespread commercial use of dialogue systems, often based on VoiceXML. Some more sophisticated systems have also seen deployment. For example Clarissa, the first spoken dialogue system used in space, is a speech-enabled procedure navigator that was used by astronauts on the International Space Station (Rayner and Hockey, 2004; Aist et al., 2002). Much research focuses on more mundane in-vehicle applications in cars Weng et al. (2006, inter alia). Among the important technical challenges in embedding these dialogue systems in real applications are good techniques for endpointing (deciding if the speaker is done talking) (Ferrer et al., 2003) and for noise robustness.

Good surveys on dialogue systems include Harris (2005), Cohen et al. (2004), McTear (2002, 2004), Sadek and De Mori (1998), and the dialogue chapter in Allen

原书第 983 页

(1995).

EXERCISES

24.1 List the dialogue act misinterpretations in the Who's On First routine at the beginning of the chapter.

24.2 Write a finite-state automaton for a dialogue manager for checking your bank balance and withdrawing money at an automated teller machine.

24.3 Dispreferred responses (for example turning down a request) are usually signaled by surface cues, such as significant silence. Try to notice the next time you or someone else utters a dispreferred response, and write down the utterance. What are some other cues in the response that a system might use to detect a dispreferred response? Consider non-verbal cues like eye-gaze and body gestures.

24.4 When asked a question to which they aren't sure they know the answer, people display their lack of confidence via cues that resemble other dispreferred responses. Try to notice some unsure answers to questions. What are some of the cues? If you have trouble doing this, read Smith and Clark (1993) and listen specifically for the cues they mention.

24.5 Build a VoiceXML dialogue system for giving the current time around the world. The system should ask the user for a city and a time format (24 hour, etc) and should return the current time, properly dealing with time zones.

24.6 Implement a small air-travel help system based on text input. Your system should get constraints from the user about a particular flight that they want to take, expressed in natural language, and display possible flights on a screen. Make simplifying assumptions. You may build in a simple flight database or you may use a flight information system on the web as your backend.

24.7 Augment your previous system to work with speech input via VoiceXML. (or alternatively, describe the user interface changes you would have to make for it to work via speech over the phone). What were the major differences?

24.8 Design a simple dialogue system for checking your email over the telephone. Implement in VoiceXML.

24.9 Test your email-reading system on some potential users. Choose some of the metrics described in Sec. 24.4.2 and evaluate your system.

原书第 984 页

Aist, G., Dowding, J., Hockey, B. A., and Hieronymus, J. L. (2002). An intelligent procedure assistant for astronaut training and support. In ACL-02, Philadelphia, PA.

Allen, J. (1995). Natural Language Understanding. Benjamin Cummings, Menlo Park, CA.

Allen, J. and Core, M. (1997). Draft of DAMSL: Dialog act markup in several layers. Unpublished manuscript.

Allen, J., Ferguson, G., and Stent, A. (2001). An architecture for more realistic conversational systems. In IUI '01: Proceedings of the 6th international conference on Intelligent user interfaces, Santa Fe, New Mexico, United States, pp. 1–8. ACM Press.

Allen, J. and Perrault, C. R. (1980). Analyzing intention in utterances. Artificial Intelligence, 15, 143–178.

Allwood, J. (1995). An activity-based approach to pragmatics. Gothenburg Papers in Theoretical Linguistics, 76.

Allwood, J., Nivre, J., and Ahlsén, E. (1992). On the semantics and pragmatics of linguistic feedback. Journal of Semantics, 9, 1–26.

Atkinson, M. and Drew, P. (1979). Order in Court. Macmillan, London.

Austin, J. L. (1962). How to Do Things with Words. Harvard University Press.

Baum, L. F. (1900). The Wizard of Oz. Available at Project Gutenberg.

Bellman, R. (1957). Dynamic Programming. Princeton University Press, Princeton, NJ.

Bobrow, D. G., Kaplan, R. M., Kay, M., Norman, D. A., Thompson, H., and Winograd, T. (1977). GUS, a frame driven dialog system. Artificial Intelligence, 8, 155–173.

Bohus, D. and Rudnicky, A. I. (2005). Sorry, I Didn't Catch That! - An Investigation of Non-understanding Errors and Recovery Strategies. In SIGdial-2005, Lisbon, Portugal.

Bouwman, G., Sturm, J., and Boves, L. (1999). Incorporating confidence measures in the Dutch train timetable information system developed in the Arise project. In IEEE ICASSP-99, pp. 493–496.

Bulyko, I., Kirchhoff, K., Ostendorf, M., and Goldberg, J. (2005). Error-sensitive response generation in a spoken langugage dialogue system. Speech Communication, 45(3), 271–288.

Bunt, H. (1994). Context and dialogue control. Think, 3, 19–31.

Bunt, H. (2000). Dynamic interpretation and dialogue theory, volume 2. In Taylor, M. M., Neel, F., and Bouwhuis, D. G. (Eds.), The structure of multimodal dialogue, pp. 139–166. John Benjamins, Amsterdam.

Carberry, S. (1990). Plan Recognition in Natural Language Dialog. MIT Press.

Carletta, J., Dahlbäck, N., Reithinger, N., and Walker, M. A. (1997a). Standards for dialogue coding in natural language processing. Tech. rep. Report no. 167, Dagstuhl Seminars. Report from Dagstuhl seminar number 9706.

Carletta, J., Isard, A., Isard, S., Kowtko, J. C., Doherty-Sneddon, G., and Anderson, A. H. (1997b). The reliability of a dialogue structure coding scheme. Computational Linguistics, 23(1), 13–32.

Chu-Carroll, J. (1998). A statistical model for discourse act recognition in dialogue interactions. In Chu-Carroll, J. and Green, N. (Eds.), Applying Machine Learning to Discourse Processing. Papers from the 1998 AAAI Spring Symposium. Tech. rep. SS-98-01, Menlo Park, CA, pp. 12–17. AAAI Press.

Chu-Carroll, J. and Brown, M. K. (1997). Tracking initiative in collaborative dialogue interactions. In ACL/EACL-97, Madrid, Spain, pp. 262–270.

Chu-Carroll, J. and Carberry, S. (1998). Collaborative response generation in planning dialogues. Computational Linguistics, 24(3), 355–400.

Chu-Carroll, J. and Carpenter, B. (1999). Vector-based natural language call routing. Computational Linguistics, 25(3), 361–388.

Chung, G. (2004). Developing a flexible spoken dialog system using simulation. In ACL-04, Barcelona, Spain.

Clark, H. H. (1994). Discourse in production. In Gernsbacher, M. A. (Ed.), Handbook of Psycholinguistics. Academic Press.

Clark, H. H. (1996). Using Language. Cambridge University Press.

Clark, H. H. and Marshall, C. (1981). Definite reference and mutual knowledge. In Joshi, A. K., Webber, B. L., and Sag, I. (Eds.), Elements of discourse understanding, pp. 10–63. Cambridge.

Clark, H. H. and Schaefer, E. F. (1989). Contributing to discourse. Cognitive Science, 13, 259–294.

Clark, H. H. and Wilkes-Gibbs, D. (1986). Referring as a collaborative process. Cognition, 22, 1–39.

Cohen, M. H., Giangola, J. P., and Balogh, J. (2004). Voice User Interface Design. Addison-Wesley, Boston.

Cohen, P. R. and Oviatt, S. L. (1994). The role of voice in human-machine communication. In Roe, D. B. and Wilpon, J. G. (Eds.), Voice Communication Between Humans and Machines, pp. 34–75. National Academy Press, Washington, D.C.

Cohen, P. R. and Perrault, C. R. (1979). Elements of a plan-based theory of speech acts. Cognitive Science, 3(3), 177–212.

Colby, K. M., Weber, S., and Hilf, F. D. (1971). Artificial para-noia. Artificial Intelligence, 2(1), 1–25.

Cole, R. A., Novick, D. G., Vermeulen, P. J. E., Sutton, S., Fanty, M., Wessels, L. F. A., de Villiers, J. H., Schalkwyk, J., Hansen, B., and Burnett, D. (1997). Experiments with a spoken dialogue system for taking the US census. Speech Communication, 23, 243–260.

Cole, R. A., Novick, D. G., Burnett, D., Hansen, B., Sutton, S., and Fanty, M. (1994). Towards automatic collection of the U.S. census. In IEEE ICASSP-94, Adelaide, Australia, Vol. I, pp. 93–96. IEEE.

原书第 985 页

Cole, R. A., Novick, D. G., Fanty, M., Sutton, S., Hansen, B., and Burnett, D. (1993). Rapid prototyping of spoken language systems: The Year 2000 Census Project. In Proceedings of the International Symposium on Spoken Dialogue, Waseda University, Tokyo, Japan.

Core, M., Ishizaki, M., Moore, J. D., Nakatani, C., Reithinger, N., Traum, D. R., and Tutiya, S. (1999). The report of the third workshop of the Discourse Resource Initiative, Chiba University and Kazusa Academia Hall. Tech. rep. No.3 CC-TR-99-1, Chiba Corpus Project, Chiba, Japan.

Cuayáhuitl, H., Renals, S., Lemon, O., and Shimodaira, H. (2007). Hierarchical dialogue optimization using semi-Markov decision processes. In INTERSPEECH-07.

Daly, N. A. and Zue, V. W. (1992). Statistical and linguistic analyses of $ F_{0} $ in read and spontaneous speech. In ICSLP-92, Vol. 1, pp. 763–766.

Danieli, M. and Gerbino, E. (1995). Metrics for evaluating dialogue strategies in a spoken language system. In Proceedings of the 1995 AAAI Spring Symposium on Empirical Methods in Discourse Interpretation and Generation, Stanford, CA, pp. 34–39. AAAI Press, Menlo Park, CA.

Ferrer, L., Shriberg, E., and Stolcke, A. (2003). A prosody-based approach to end-of-utterance detection that does not re-quire speech recognition. In IEEE ICASSP-03.

Fikes, R. E. and Nilsson, N. J. (1971). STRIPS: A new approach to the application of theorem proving to problem solving. Artificial Intelligence, 2, 189–208.

Fraser, N. (1992). Assessment of interactive systems. In Gibbon, D., Moore, R., and Winski, R. (Eds.), Handbook on Standards and Resources for Spoken Language Systems, pp. 564–615. Mouton de Gruyter, Berlin.

Fraser, N. M. and Gilbert, G. N. (1991). Simulating speech systems. Computer Speech and Language, 5, 81–99.

Good, M. D., Whiteside, J. A., Wixon, D. R., and Jones, S. J. (1984). Building a user-derived interface. Communications of the ACM, 27(10), 1032–1043.

Goodwin, C. (1996). Transparent vision. In Ochs, E., Schegloff, E. A., and Thompson, S. A. (Eds.), Interaction and Grammar, pp. 370–404. Cambridge University Press.

Gordon, D. and Lakoff, G. (1971). Conversational postulates. In CLS-71, pp. 200–213. University of Chicago. Reprinted in Peter Cole and Jerry L. Morgan (Eds.), Speech Acts: Syntax and Semantics Volume 3, Academic, 1975.

Gorin, A. L., Riccardi, G., and Wright, J. H. (1997). How may i help you?. Speech Communication, 23, 113–127.

Gould, J. D., Conti, J., and Hovanyecz, T. (1983). Composing letters with a simulated listening typewriter. Communications of the ACM, 26(4), 295–308.

Gould, J. D. and Lewis, C. (1985). Designing for usability: Key principles and what designers think. Communications of the ACM, 28(3), 300–311.

Grice, H. P. (1957). Meaning. Philosophical Review, 67, 377–388. Reprinted in Semantics, edited by Danny D. Steinberg

& Leon A. Jakobovits (1971), Cambridge University Press, pages 53–59.

Grice, H. P. (1975). Logic and conversation. In Cole, P. and Morgan, J. L. (Eds.), Speech Acts: Syntax and Semantics Volume 3, pp. 41–58. Academic Press.

Grice, H. P. (1978). Further notes on logic and conversation. In Cole, P. (Ed.), Pragmatics: Syntax and Semantics Volume 9, pp. 113–127. Academic Press.

Grosz, B. J. and Hirschberg, J. (1992). Some intonational characteristics of discourse structure. In ICSLP-92, Vol. 1, pp. 429–432.

Grosz, B. J. (1977). The Representation and Use of Focus in Dialogue Understanding. Ph.D. thesis, University of California, Berkeley.

Grosz, B. J. and Sidner, C. L. (1980). Plans for discourse. In Cohen, P. R., Morgan, J., and Pollack, M. E. (Eds.), Intentions in Communication, pp. 417–444. MIT Press.

Grosz, B. J. and Sidner, C. L. (1986). Attention, intentions, and the structure of discourse. Computational Linguistics, 12(3), 175–204.

Guindon, R. (1988). A multidisciplinary perspective on dialogue structure in user-advisor dialogues. In Guindon, R. (Ed.), Cognitive Science And Its Applications For Human-Computer Interaction, pp. 163–200. Lawrence Erlbaum.

Harris, R. A. (2005). Voice Interaction Design: Crafting the New Conversational Speech Systems. Morgan Kaufmann.

Hemphill, C. T., Godfrey, J., and Doddington, G. (1990). The ATIS spoken language systems pilot corpus. In Proceedings DARPA Speech and Natural Language Workshop, Hidden Valley, PA, pp. 96–101. Morgan Kaufmann.

Hinkelman, E. A. and Allen, J. (1989). Two constraints on speech act ambiguity. In Proceedings of the 27th ACL, Vancouver, Canada, pp. 212–219.

Hintikka, J. (1969). Semantics for propositional attitudes. In Davis, J. W., Hockney, D. J., and Wilson, W. K. (Eds.), Philosophical Logic, pp. 21–45. D. Reidel, Dordrecht, Holland.

Hirschberg, J. and Litman, D. J. (1993). Empirical studies on the disambiguation of cue phrases. Computational Linguistics, 19(3), 501–530.

Hirschberg, J., Litman, D. J., and Swerts, M. (2001). Identifying user corrections automatically in spoken dialogue systems. In NAACL.

Hirschberg, J. and Nakatani, C. (1996). A prosodic analysis of discourse segments in direction-giving monologues. In ACL-96, Santa Cruz, CA, pp. 286–293.

Hirschberg, J. and Pierrehumbert, J. B. (1986). The intonational structuring of discourse. In ACL-86, New York, pp. 136–144.

Hirschman, L. and Pao, C. (1993). The cost of errors in a spoken language system. In EUROSPEECH-93, pp. 1419–1422.

Issar, S. and Ward, W. (1993). Cmu's robust spoken language understanding system. In Eurospeech 93, pp. 2147–2150.

Jefferson, G. (1984). Notes on a systematic deployment of the acknowledgement tokens 'yeah' and 'mm hm'. Papers in Linguistics, 17(2), 197–216.

原书第 986 页

Jekat, S., Klein, A., Maier, E., Maleck, I., Mast, M., and Quantz, J. (1995). Dialogue Acts in VERBMOBIL verbmobil-report-65-95..

Johnston, M., Ehlen, P., Gibbon, D., and Liu, Z. (2007). The multimodal presentation dashboard. In NAACL HLT 2007 Workshop 'Bridging the Gap'.

Kamm, C. A. (1994). User interfaces for voice applications. In Roe, D. B. and Wilpon, J. G. (Eds.), Voice Communication Between Humans and Machines, pp. 422–442. National Academy Press, Washington, D.C.

Kita, K., Fukui, Y., Nagata, M., and Morimoto, T. (1996). Automatic acquisition of probabilistic dialogue models. In ICSLP-96, Philadelphia, PA, Vol. 1, pp. 196–199.

Kompe, R., Kießling, A., Kuhn, T., Mast, M., Niemann, H., Nöth, E., Ott, K., and Batliner, A. (1993). Prosody takes over: A prosodically guided dialog system. In EUROSPEECH-93, Berlin, Vol. 3, pp. 2003–2006.

Labov, W. and Fanshel, D. (1977). Therapeutic Discourse. Academic Press.

Landauer, T. K. (Ed.). (1995). The Trouble With Computers: Usefulness, Usability, and Productivity. MIT Press.

Lemon, O., Georgila, K., Henderson, J., and Stuttle, M. (2006). An ISU dialogue system exhibiting reinforcement learning of dialogue policies: generic slot-filling in the TALK in-car system. In EACL-06.

Levesque, H. J., Cohen, P. R., and Nunes, J. H. T. (1990). On acting together. In AAAI-90, Boston, MA, pp. 94–99. Morgan Kaufmann.

Levin, E., Pieraccini, R., and Eckert, W. (2000). A stochastic model of human-machine interaction for learning dialog strategies. IEEE Transactions on Speech and Audio Processing, 8, 11–23.

Levinson, S. C. (1983). Pragmatics. Cambridge University Press.

Levow, G.-A. (1998). Characterizing and recognizing spoken corrections in human-computer dialogue. In COLING-ACL, pp. 736–742.

Lewin, I., Becket, R., Boye, J., Carter, D., Rayner, M., and Wirén, M. (1999). Language processing for spoken dialogue systems: is shallow parsing enough?. In Accessing Information in Spoken Audio: Proceedings of ESCA ETRW Workshop, Cambridge, 19 & 20th April 1999, pp. 37–42.

Litman, D. J. (1985). Plan Recognition and Discourse Analysis: An Integrated Approach for Understanding Dialogues. Ph.D. thesis, University of Rochester, Rochester, NY.

Litman, D. J. and Allen, J. (1987). A plan recognition model for subdialogues in conversation. Cognitive Science, 11, 163–200.

Litman, D. J. and Pan, S. (2002). Designing and evaluating an adaptive spoken dialogue system. User Modeling and User-Adapted Interaction, 12(2-3), 111–137.

Litman, D. J. and Silliman, S. (2004). Itspoke: An intelligent tutoring spoken dialogue system. In HLT-NAACL-04.

Litman, D. J., Swerts, M., and Hirschberg, J. (2000). Predicting automatic speech recognition performance using prosodic cues. In NAACL 2000.

Litman, D. J., Walker, M. A., and Kearns, M. S. (1999). Automatic detection of poor speech recognition at the dialogue level. In ACL-99, College Park, MA, pp. 309–316. ACL.

Lochbaum, K. E. (1998). A collaborative planning model of intentional structure. Computational Linguistics, 24(4), 525–572.

Lochbaum, K. E., Grosz, B. J., and Sidner, C. L. (1990). Models of plans to support communication: An initial report. In AAAI-90, Boston, MA, pp. 485–490. Morgan Kaufmann.

Lochbaum, K. E., Grosz, B. J., and Sidner, C. L. (2000). Discourse structure and intention recognition. In Dale, R., Somers, H. L., and Moisl, H. (Eds.), Handbook of Natural Language Processing. Marcel Dekker.

McTear, M. F. (2002). Spoken dialogue technology: Enabling the conversational interface. ACM Computing Surveys, 34(1), 90–169.

McTear, M. F. (2004). Spoken Dialogue Technology. Springer Verlag, London.

Miller, S., Bobrow, R. J., Ingria, R., and Schwartz, R. (1994). Hidden understanding models of natural language. In Proceedings of the 32nd ACL, Las Cruces, NM, pp. 25–32.

Miller, S., Fox, H., Ramshaw, L. A., and Weischedel, R. (2000). A novel use of statistical parsing to extract information from text. In Proceedings of the 1st Annual Meeting of the North American Chapter of the ACL (NAACL), Seattle, Washington, pp. 226–233.

Miller, S., Stallard, D., Bobrow, R. J., and Schwartz, R. (1996). A fully statistical approach to natural language interfaces. In ACL-96, Santa Cruz, CA, pp. 55–61.

Möller, S. (2002). A new taxonomy for the quality of telephone services based on spoken dialogue systems. In In Proceedings of the 3rd SIGdial Workshop on Discourse and Dialogue, pp. 142–153.

Möller, S. (2004). Quality of Telephone-Based Spoken Dialogue Systems. Springer.

Nagata, M. and Morimoto, T. (1994). First steps toward statistical modeling of dialogue to predict the speech act type of the next utterance. Speech Communication, 15, 193–203.

Niekrasz, J. and Purver, M. (2006). A multimodal discourse ontology for meeting understanding. In Renals, S. and Bengio, S. (Eds.), Machine Learning for Multimodal Interaction: Second International Workshop MLMI 2005, Revised Selected Papers, No. 3689 in Lecture Notes in Computer Science, pp. 162–173. Springer-Verlag.

Nielsen, J. (1992). The usability engineering life cycle. IEEE Computer, 25(3), 12–22.

Norman, D. A. (1988). The Design of Everyday Things. Basic Books.

Oviatt, S. L., Cohen, P. R., Wang, M. Q., and Gaston, J. (1993). A simulation-based research strategy for designing complex

原书第 987 页

NL systems. In Proceedings DARPA Speech and Natural Language Workshop, Princeton, NJ, pp. 370–375. Morgan Kaufmann.

Oviatt, S. L., MacEachern, M., and Levow, G.-A. (1998). Predicting hyperarticulate speech during human-computer error resolution. Speech Communication, 24, 87–110.

Passonneau, R. and Litman, D. J. (1993). Intention-based segmentation: Human reliability and correlation with linguistic cues. In Proceedings of the 31st ACL, Columbus, Ohio, pp. 148–155.

Perrault, C. R. and Allen, J. (1980). A plan-based analysis of indirect speech acts. American Journal of Computational Linguistics, 6(3-4), 167–182.

Pieraccini, R., Levin, E., and Lee, C.-H. (1991). Stochastic representation of conceptual structure in the ATIS task. In Proceedings DARPA Speech and Natural Language Workshop, Pacific Grove, CA, pp. 121–124. Morgan Kaufmann.

Pierrehumbert, J. B. and Hirschberg, J. (1990). The meaning of intonational contours in the interpretation of discourse. In Cohen, P. R., Morgan, J., and Pollack, M. (Eds.), Intentions in Communication, pp. 271–311. MIT Press.

Pierrehumbert, J. B. (1980). The Phonology and Phonetics of English Intonation. Ph.D. thesis, MIT.

Polifroni, J., Hirschman, L., Seneff, S., and Zue, V. W. (1992). Experiments in evaluating interactive spoken language systems. In Proceedings DARPA Speech and Natural Language Workshop, Harriman, NY, pp. 28–33. Morgan Kaufmann.

Power, R. (1979). The organization of purposeful dialogs. Linguistics, 17, 105–152.

Rayner, M. and Hockey, B. A. (2003). Transparent combination of rule-based and data-driven approaches in a speech understanding architecture. In EACL-03, Budapest, Hungary.

Rayner, M. and Hockey, B. A. (2004). Side effect free dialogue management in a voice enabled procedure browser. In ICSLP-04, pp. 2833–2836.

Rayner, M., Hockey, B. A., and Bouillon, P. (2006). Putting Linguistics into Speech Recognition. CSLI.

Reichman, R. (1985). Getting Computers to Talk Like You and Me. MIT Press.

Reiter, E. and Dale, R. (2000). Building Natural Language Generation Systems. Cambridge University Press.

Reithinger, N., Engel, R., Kipp, M., and Klesen, M. (1996). Predicting dialogue acts for a speech-to-speech translation system. In ICSLP-96, Philadelphia, PA, Vol. 2, pp. 654–657.

Roy, N., Pineau, J., and Thrun, S. (2000). Spoken dialog management for robots. In ACL-00, Hong Kong.

Russell, S. and Norvig, P. (2002). Artificial Intelligence: A Modern Approach. Prentice Hall. Second edition.

Sacks, H., Schegloff, E. A., and Jefferson, G. (1974). A simplest systematics for the organization of turn-taking for conversation. Language, 50(4), 696–735.

Sadek, D. and De Mori, R. (1998). Dialogue systems. In De Mori, R. (Ed.), Spoken Dialogues With Computers. Academic Press, London.

Sadek, M. D. (1991). Dialogue acts are rational plans. In ESCA/ETR Workshop on the Structure of Multimodal Dialogue, pp. 19–48.

Sag, I. A. and Liberman, M. Y. (1975). The intonational dis-

ambiguation of indirect speech acts. In CLS-75, pp. 487–498.

University of Chicago.

San-Segundo, R., Montero, J. M., Ferreiros, J., Córdoba, R., and Pardo, J. M. (2001). Designing confirmation mechanisms and error recovery techniques in a railway information system for Spanish. In In Proceedings of the 2nd SIGdial Workshop on Discourse and Dialogue, Aalborg, Denmark.

Schegloff, E. A. (1968). Sequencing in conversational openings. American Anthropologist, 70, 1075–1095.

Schegloff, E. A. (1979). Identification and recognition in telephone conversation openings. In Psathas, G. (Ed.), Everyday Language: Studies in Ethnomethodology, pp. 23–78. Irvington.

Schegloff, E. A. (1982). Discourse as an interactional achievement: Some uses of ‘uh huh’ and other things that come between sentences. In Tannen, D. (Ed.), Analyzing Discourse: Text and Talk, pp. 71–93. Georgetown University Press, Washington, D.C.

Searle, J. R. (1975a). Indirect speech acts. In Cole, P. and Morgan, J. L. (Eds.), Speech Acts: Syntax and Semantics Volume 3, pp. 59–82. Academic Press.

Searle, J. R. (1975b). A taxonomy of illocutionary acts. In Gunderson, K. (Ed.), Language, Mind and Knowledge, Minnesota Studies in the Philosophy of Science, Vol. VII, pp. 344–369. University of Minnesota Press, Amsterdam. Also appears in John R. Searle, Expression and Meaning: Studies in the Theory of Speech Acts, Cambridge University Press, 1979.

Seneff, S. (1995). TINA: A natural language system for spoken language application. Computational Linguistics, 18(1), 62–86.

Seneff, S. (2002). Response planning and generation in the MERCURY flight reservation system. Computer Speech and Language, Special Issue on Spoken Language Generation, 16(3-4), 283–312.

Seneff, S. and Polifroni, J. (2000). Dialogue management in the mercury flight reservation system. In ANLP/NAACL Workshop on Conversational Systems, Seattle.

Shriberg, E., Bates, R., Taylor, P., Stolcke, A., Jurafsky, D., Ries, K., Coccaro, N., Martin, R., Meteer, M., and Van Ess-Dykema, C. (1998). Can prosody aid the automatic classification of dialog acts in conversational speech?. Language and Speech (Special Issue on Prosody and Conversation), 41(3-4), 439–487.

Shriberg, E., Wade, E., and Price, P. (1992). Human-machine problem solving using spoken language systems (SLS): Factors affecting performance and user satisfaction. In Proceedings DARPA Speech and Natural Language Workshop, Hariman, NY, pp. 49–54. Morgan Kaufmann.

原书第 988 页

Singh, S. P., Litman, D. J., Kearns, M. J., and Walker, M. A. (2002). Optimizing dialogue management with reinforcement learning: Experiments with the njfun system. J. Artif. Intell. Res. (JAIR), 16, 105–133.

Smith, R. W. and Gordon, S. A. (1997). Effects of variable initiative on linguistic behavior in human-computer spoken natural language dialogue. Computational Linguistics, 23(1), 141–168.

Smith, V. L. and Clark, H. H. (1993). On the course of answering questions. Journal of Memory and Language, 32, 25–38.

Stalnaker, R. C. (1978). Assertion. In Cole, P. (Ed.), Pragmatics: Syntax and Semantics Volume 9, pp. 315–332. Academic Press.

Stent, A. (2002). A conversation acts model for generating spoken dialogue contributions. Computer Speech and Language, Special Issue on Spoken Language Generation, 16(3-4).

Stifelman, L. J., Arons, B., Schmandt, C., and Hulteen, E. A. (1993). VoiceNotes: A speech interface for a hand-held voice notetaker. In Human Factors in Computing Systems: INTER-CHI '93 Conference Proceedings, Amsterdam, pp. 179–186. ACM.

Stolcke, A., Ries, K., Coccaro, N., Shriberg, E., Bates, R., Jurafsky, D., Taylor, P., Martina, R., Meteer, M., and Van Ess-Dykema, C. (2000). Dialogue act modeling for automatic tagging and recognition of conversational speech. Computational Linguistics, 26, 339–371.

Stolcke, A., Shriberg, E., Bates, R., Coccaro, N., Jurafsky, D., Martin, R., Meteer, M., Ries, K., Taylor, P., and Van Ess-Dykema, C. (1998). Dialog act modeling for conversational speech. In Chu-Carroll, J. and Green, N. (Eds.), Applying Machine Learning to Discourse Processing. Papers from the 1998 AAAI Spring Symposium. Tech. rep. SS-98-01, Stanford, CA, pp. 98–105. AAAI Press.

Sutton, R. S. and Barto, A. G. (1998). Reinforcement Learning: An Introduction. Bradford Books (MIT Press).

Swerts, M., Litman, D. J., and Hirschberg, J. (2000). Corrections in spoken dialogue systems. In ICSLP-00, Beijing, China.

Taylor, P., King, S., Isard, S., and Wright, H. (1998). Intonation and dialog context as constraints for speech recognition. Language and Speech, 41(3-4), 489–508.

Traum, D. R. (2000). 20 questions for dialogue act taxonomies. Journal of Semantics, 17(1).

Traum, D. R. and Hinkelman, E. A. (1992). Conversation acts in task-oriented spoken dialogue. Computational Intelligence: Special Issue on Computational Approaches to Non-Literal Language, 8(3).

Traum, D. R. and Larsson, S. (2000). Information state and dialogue management in the trindi dialogue move engine toolkit. Natural Language Engineering, 6(323-340), 97–114.

Traum, D. R. and Larsson, S. (2003). The information state approach to dialogue management. In van Kuppevelt, J. and Smith, R. (Eds.), Current and New Directions in Discourse and Dialogue. Kluwer.

VanLehn, K., Jordan, P. W., Rosé, C., Bhembe, D., Böttner, M., Gaydos, A., Makatchev, M., Pappuswamy, U., Ringenberg, M., Roque, A., Siler, S., Srivastava, R., and Wilson, R. (2002). The architecture of Why2-Atlas: A coach for qualitative physics essay writing. In Proc. Intelligent Tutoring Systems.

Wade, E., Shriberg, E., and Price, P. J. (1992). User behaviors affecting speech recognition. In ICSLP-92, pp. 995–998.

Waibel, A. (1988). Prosody and Speech Recognition. Morgan Kaufmann.

Walker, M. A., Fromer, J. C., and Narayanan, S. S. (1998). Learning optimal dialogue strategies: a case study of a spoken dialogue agent for email. In COLING/ACL-98, Montreal, Canada, pp. 1345–1351.

Walker, M. A., Kamm, C. A., and Litman, D. J. (2001). Towards developing general models of usability with PARADISE. Natural Language Engineering: Special Issue on Best Practice in Spoken Dialogue Systems, 6(3).

Walker, M. A., Litman, D. J., Kamm, C. A., and Abella, A. (1997). PARADISE: A framework for evaluating spoken dialogue agents. In ACL/EACL-97, Madrid, Spain, pp. 271–280.

Walker, M. A., Maier, E., Allen, J., Carletta, J., Condon, S., Flammia, G., Hirschberg, J., Isard, S., Ishizaki, M., Levin, L., Luperfoy, S., Traum, D. R., and Whittaker, S. (1996). Penn multiparty standard coding scheme: Draft annotation manual. www.cis.upenn.edu/~ircs/dis course-tagging/newcoding.html.

Walker, M. A., Passonneau, R., Rudnicky, A. I., Aberdeen, J., Boland, J., Bratt, E., Garofolo, J., Hirschman, L., Le, A., Lee, S., Narayanan, S. S., Papineni, K., Pellom, B., Polifroni, J., Potamianos, A., Prabhu, P., Rudnicky, A. I., Sanders, G., Seneff, S., Stallard, D., and Whittaker, S. (2002). Cross-site evaluation in DARPA Communicator: The June 2000 data collection submitted.

Walker, M. A. and Rambow, O. (2002). Spoken language generation. Computer Speech and Language, Special Issue on Spoken Language Generation, 16(3-4), 273–281.

Walker, M. A. and Whittaker, S. (1990). Mixed initiative in dialogue: An investigation into discourse segmentation. In Proceedings of the 28th ACL, Pittsburgh, PA, pp. 70–78.

Walker, M. A. et al. (2001). Cross-site evaluation in darpa communicator: The June 2000 data collection. Submitted ms.

Ward, N. and Tsukahara, W. (2000). Prosodic features which cue back-channel feedback in English and Japanese. Journal of Pragmatics, 32, 1177–1207.

Ward, W. and Issar, S. (1994). Recent improvements in the cmu spoken language understanding system. In ARPA Human Language Technologies Workshop, Plainsboro, N.J.

Warnke, V., Kompe, R., Niemann, H., and Nöth, E. (1997). Integrated dialog act segmentation and classification using prosodic features and language models. In EUROSPEECH97, Vol. 1, pp. 207–210.

Weinschenk, S. and Barker, D. T. (2000). Designing effective speech interfaces. Wiley.

原书第 989 页

Weng, F., Varges, S., Raghunathan, B., Ratiu, F., Pon-Barry, H., Lathrop, B., Zhang, Q., Scheideck, T., Bratt, H., Xu, K., Purver, M., Mishra, R., Raya, M., Peters, S., Meng, Y., Cavedon, L., and Shriberg, E. (2006). Chat: A conversational helper for automotive tasks. In ICSLP-06, pp. 1061–1064.

Wilensky, R. (1983). Planning and Understanding: A Computational Approach to Human Reasoning. Addison-Wesley.

Williams, J. D. and Young, S. J. (2000). Partially observable markov decision processes for spoken dialog systems. Computer Speech and Language, 21(1), 393–422.

Williams, J. D. and Young, S. J. (2005). Scaling up pomdps for dialog management: The "summary pomdp" method. In IEEE ASRU-05.

Wittgenstein, L. (1953). Philosophical Investigations. (Translated by Anscombe, G.E.M.). Blackwell, Oxford.

Woszczyna, M. and Waibel, A. (1994). Inferring linguistic structure in spoken language. In ICSLP-94, Yokohama, Japan, pp. 847–850.

Xu, W. and Rudnicky, A. I. (2000). Task-based dialog management using an agenda. In ANLP/NAACL Workshop on Conversational Systems, Somerset, New Jersey, pp. 42–47.

Yankelovich, N., Levow, G.-A., and Marx, M. (1995). Designing SpeechActs: Issues in speech user interfaces. In Human Factors in Computing Systems: CHI '95 Conference Proceedings, Denver, CO, pp. 369–376. ACM.

Yngve, V. H. (1970). On getting a word in edgewise. In CLS-70, pp. 567–577. University of Chicago.

Young, S. J. (2002). The statistical approach to the design of spoken dialogue systems. Tech. rep. CUED/F-INFENG/TR.433, Cambridge University Engineering Department, Cambridge, England.

Zue, V. W., Glass, J., Goodine, D., Leung, H., Phillips, M., Polifroni, J., and Seneff, S. (1989). Preliminary evaluation of the VOYAGER spoken language system. In Proceedings DARPA Speech and Natural Language Workshop, Cape Cod, MA, pp. 160–167. Morgan Kaufmann.

原书第 990 页

25

The process of translating comprises in its essence the whole secret of human understanding and social communication.

Attributed to Hans-Georg Gadamer

What is translation? On a platter

A poet's pale and glaring head,

A parrot's screech, a monkey's chatter,

And profanation of the dead.

Nabokov, On Translating Eugene Onegin

Proper words in proper places

Jonathan Swift

MACHINE TRANSLATION MT

This chapter introduces techniques for machine translation (MT), the use of computers to automate some or all of the process of translating from one language to another. Translation, in its full generality, is a difficult, fascinating, and intensely human endeavor, as rich as any other area of human creativity. Consider the following passage from the end of Chapter 45 of the 18th-century novel The Story of the Stone, also called Dream of the Red Chamber, by Cao Xue Qin (Cao, 1792), transcribed in the Mandarin dialect:

dai yu zi zai chuang shang gan nian bao chai... you ting jian chuang wai zhu shao xiang ye zhe shang, yu sheng xi li, qing han tou mu, bu jue you di xia lei lai.

Fig. 25.1 shows the English translation of this passage by David Hawkes, in sentences labeled $ E_{1} $- $ E_{4} $. For ease of reading, instead of giving the Chinese, we have shown the English glosses of each Chinese word IN SMALL CAPS. Words in blue are Chinese words not translated into English, or English words not in the Chinese. We have shown alignment lines between words that roughly correspond in the two languages.

Consider some of the issues involved in this translation. First, the English and Chinese texts are very different structurally and lexically. The four English sentences

原书第 991 页
Image
Figure 25.1 A Chinese passage from Dream of the Red Chamber, with the Chinese words represented by English glosses IN SMALL CAPS. Alignment lines are drawn between ‘Chinese’ words and their English translations. Words in italics are Chinese words not translated into English, or English words not in the original Chinese.

(notice the periods in blue) correspond to one long Chinese sentence. The word order of the two texts is very different, as we can see by the many crossed alignment lines in Fig. 25.1. The English has many more words than the Chinese, as we can see by the large number of English words marked in blue. Many of these differences are caused by structural differences between the two languages. For example, because Chinese rarely marks verbal aspect or tense; the English translation has additional words like as, turned to, and had begun, and Hawkes had to decide to translate Chinese to as penetrated, rather than say was penetrating or had penetrated. Chinese has less articles than English, explaining the large number of blue thes. Chinese also uses far fewer pronouns than English, so Hawkes had to insert she and her in many places into the English translation.

Stylistic and cultural differences are another source of difficulty for the translator. Unlike English names, Chinese names are made up of regular content words with meanings. Hawkes chose to use transliterations (Daiyu) for the names of the main characters but to translate names of servants by their meanings (Aroma, Skybright). To make the image clear for English readers unfamiliar with Chinese bed-curtains, Hawkes translated ma ('curtain') as curtains of her bed. The phrase bamboo tip plantain leaf, although elegant in Chinese, where such four-character phrases are a hallmark of literate prose, would be awkward if translated word-for-word into English, and so Hawkes used simply bamboos and plantains.

Translation of this sort clearly requires a deep and rich understanding of the source language and the input text, and a sophisticated, poetic, and creative command of the

原书第 992 页

target language. The problem of automatically performing high-quality literary translation between languages as different as Chinese to English is thus far too hard to automate completely.

However, even non-literary translations between such similar languages as English and French can be difficult. Here is an English sentence from the Hansards corpus of Canadian parliamentary proceedings, with its French translation:

English: Following a two-year transitional period, the new Foodstuffs Ordinance for Mineral Water came into effect on April 1, 1988. Specifically, it contains more stringent requirements regarding quality consistency and purity guarantees.

French: La nouvelle ordonnance fédérale sur les denrées alimentaires concernant entre autres les eaux minérales, entrée en vigueur le 1er avril 1988 après une période transitoire de deux ans. exige surtout une plus grande constance dans la qualité et une garantie de la pureté.

French gloss: THE NEW ORDINANCE FEDERAL ON THE STUFF FOOD CONCERNING AMONG OTHERS THE WATERS MINERAL CAME INTO EFFECT THE 1ST APRIL 1988 AFTER A PERIOD TRANSITORY OF TWO YEARS REQUIRES ABOVE ALL A LARGER CONSISTENCY IN THE QUALITY AND A GUARANTEE OF THE PURITY.

Despite the strong structural and vocabulary overlaps between English and French, such translation, like literary translation, still has to deal with differences in word order (e.g., the location of the following a two-year transitional period phrase) and in structure (e.g., English uses the noun requirements while the French uses the verb exige 'REQUIRE').

Nonetheless, such translations are much easier, and a number of non-literary translation tasks can be addressed with current computational models of machine translation, including: (1) tasks for which a rough translation is adequate, (2) tasks where a human post-editor is used, and (3) tasks limited to small sublanguage domains in which fully automatic high quality translation (FAHQT) is still achievable.

Information acquisition on the web is the kind of task where a rough translation may still be useful. Suppose you were at the market this morning and saw some lovely plátanos (plantains, a kind of banana) at the local Caribbean grocery store and you want to know how to cook them. You go to the web, and find the following recipe:

Para 6 personas

3 Plátanos maduros

2 cucharadas de mantequilla derretida

1 taza de jugo (zumo) de naranja

5 cucharadas de azúcar morena o blanc

Pelar los plátanos, cortarlos por la mitad y, luego, a lo largo. Engrasar una fuente o pirex con margarina. Colocar los plátanos y bañarlos con la mantequilla derretida. En un recipiente hondo, mezclar el jugo (zumo) de naranja con el azúcar, jengibre, nuez moscada y ralladura de naranja. Verter sobre los plátanos y hornear a $ 325^{\circ} $ F. Los primeros 15 minutos, dejar los pátanos cubiertos, hornear 10 o 15 minutos más destapando los plátanos

原书第 993 页
Platano in OrangeFor 6 people
3 Bananas mature2 tablespoon melted butter
1 cup juice (juice) orange5 tablespoons brown sugar or white
1/8 teaspoon nutmeg powder1 tablespoon ralladura orange
1 tablespoon cinnamon powder (optional)
Peel bananas, cut in half and then along. Grease a source or pirex with margarine. Put bananas and showering them with the melted butter. In a deep bowl, mix the juice (juice) orange with the sugar, ginger, nutmeg and ralladura orange. Pour over bananas and bake to 350° F. The first 15 minutes, leave covered bananas, bake 10 to 15 minutes more uncovering bananas.

While there are still lots of confusions in this translation (is it for bananas or plantains? What exactly is the pot we should use? What is ralladura?) it's probably enough, perhaps after looking up one or two words, to get a basic idea of something to try in the kitchen with your new purchase!

An MT system can also be used to speed-up the human translation process, by producing a draft translation that is fixed up in a post-editing phase by a human translator. Strictly speaking, systems used in this way are doing computer-aided human translation (CAHT or CAT) rather than (fully automatic) machine translation. This model of MT usage is effective especially for high volume jobs and those requiring quick turn-around, such as the translation of software manuals for localization to reach new markets.

Weather forecasting is an example of a sublanguage domain that can be modeled completely enough to use raw MT output even without post-editing. Weather forecasts consist of phrases like Cloudy with a chance of showers today and Thursday, or Outlook for Friday: Sunny. This domain has a limited vocabulary and only a few basic phrase types. Ambiguity is rare, and the senses of ambiguous words are easily disambiguated based on local context, using word classes and semantic features such as WEEKDAY, PLACE, or TIME POINT. Other domains that are sublanguage-like include equipment maintenance manuals, air travel queries, appointment scheduling, and restaurant recommendations.

Applications for machine translation can also be characterized by the number and direction of the translations. Localization tasks like translations of computer manuals require one-to-many translation (from English into many languages). One-to-many translation is also needed for non-English speakers around the world to access web information in English. Conversely, many-to-one translation (into English) is relevant for anglophone readers who need the gist of web content written in other languages. Many-to-many translation is relevant for environments like the European Union, where 23 official languages (at the time of this writing) need to be intertranslated.

Before we turn to MT systems, we begin in section 25.1 by summarizing key differences among languages. The three classic models for doing MT are then presented in Sec. 25.2: the direct, transfer, and interlingua approaches. We then investigate in detail modern statistical MT in Secs. 25.3-25.8, finishing in Sec. 25.9 with a discussion of evaluation.

原书第 994 页

25.1 WHY IS MACHINE TRANSLATION SO HARD?

We began this chapter with some of the issues that made it hard to translate The Story of the Stone from Chinese to English. In this section we look in more detail about what makes translation difficult. We'll discuss what makes languages similar or different, including systematic differences that we can model in a general way, as well as idiosyncratic and lexical differences that must be dealt with one by one. These differences between languages are referred to as translation divergences and an understanding of what causes them will help us in building models that overcome the differences (Dorr, 1994).

← 24.7.1 Plan-Inferential Interpretation and Production25.1.1 Typology →