← 学习库 Speech and Language Processing 本册目录

21.6 Consider passages (21.100a-b), adapted from Winograd (1972)

The city council denied the demonstrators a permit because

b. they advocated violence.

a. they feared violence.

What are the correct interpretations for the pronouns in each case? Sketch out an analysis of each in the interpretation as abduction framework, in which these reference assignments are made as a by-product of establishing the Explanation relation.

21.7 Select an editorial column from your favorite newspaper, and determine the discourse structure for a 10-20 sentence portion. What problems did you encounter? Were you helped by superficial cues the speaker included (e.g., discourse connectives) in any places?

原书第 827 页

Alshawi, H. (1987). Memory and Context for Language Interpretation. Cambridge University Press.

Aone, C. and Bennett, S. W. (1995). Evaluating automated and manual acquisition of anaphora resolution strategies. In ACL-95, Cambridge, MA, pp. 122–129.

Ariel, M. (2001). Accessibility theory: An overview. In Sanders, T., Schilperoord, J., and Spooren, W. (Eds.), Text Representation: Linguistic and Psycholinguistic Aspects, pp. 29–87. Benjamins.

Ariel, M. (1990). Accessing Noun Phrase Antecedents. Routledge.

Asher, N. (1993). Reference to Abstract Objects in Discourse. SLAP 50, Dordrecht, Kluwer.

Asher, N. and Lascarides, A. (2003). Logics of Conversation. Cambridge University Press.

Baldridge, J., Asher, N., and Hunter, J. (2007). Annotation for and robust parsing of discourse structure on unrestricted texts. Zeitschrift für Sprachwissenschaft. In press.

Barzilay, R. and Lapata, M. (2007). Modeling local coherence: an entity-based approach. Computational Linguistics. To appear.

Bean, D. and Riloff, E. (1999). Corpus-based identification of non-anaphoric noun phrases. In ACL-99, pp. 373–380.

Bean, D. and Riloff, E. (2004). Unsupervised learning of contextual role knowledge for coreference resolution. In HLTIAAACL-04.

Beeferman, D., Berger, A., and Lafferty, J. D. (1999). Statistical Models for Text Segmentation. Machine Learning, 34(1), 177–210.

Bergsma, S. and Lin, D. (2006). Bootstrapping path-based pronoun resolution. In COLING/ACL 2006, Sydney, Australia, pp. 33–40.

Bestgen, Y. (2006). Improving Text Segmentation Using Latent Semantic Analysis: A Reanalysis of Choi, Wiemer-Hastings, and Moore (2001). Computational Linguistics, 32(1), 5–12.

Brants, T., Chen, F., and Tsochantaridis, I. (2002). Topic-based document segmentation with probabilistic latent semantic analysis. In CIKM '02: Proceedings of the eleventh international conference on Information and knowledge management, pp. 211–218.

Brennan, S. E. (1995). Centering attention in discourse. Language and Cognitive Processes, 10, 137–167.

Brennan, S. E., Friedman, M. W., and Pollard, C. (1987). A centering approach to pronouns. In ACL-87, Stanford, CA, pp. 155–162.

Caramazza, A., Grober, E., Garvey, C., and Yates, J. (1977). Comprehension of anaphoric pronouns. Journal of Verbal Learning and Verbal Behaviour, 16, 601–609.

Cardie, C. and Wagstaff, K. (1999). Noun phrase coreference as clustering. In EMNLP/VLC-99, College Park, MD.

Carlson, L. and Marcu, D. (2001). Discourse tagging manual. Tech. rep. ISI-TR-545, ISI.

Carlson, L., Marcu, D., and Okurowski, M. E. (2001). Building a discourse-tagged corpus in the framework of rhetorical structure theory. In Proceedings of SIGDIAL.

Carlson, L., Marcu, D., and Okurowski, M. E. (2002). Building a discourse-tagged corpus in the framework of rhetorical structure theory. In van Kuppevelt, J. and Smith, R. (Eds.), Current Directions in Discourse and Dialogue. Kluwer.

Chafe, W. L. (1976). Givenness, contrastiveness, definiteness, subjects, topics, and point of view. In Li, C. N. (Ed.), Subject and Topic, pp. 25–55. Academic Press.

Charniak, E. and Shimony, S. E. (1990). Probabilistic semantics for cost based abduction. In Dietterich, T. S. W. (Ed.), AAAI-90, pp. 106–111. MIT Press.

Charniak, E. and Goldman, R. (1988). A logic for semantic interpretation. In Proceedings of the 26th ACL, Buffalo, NY.

Charniak, E. and McDermott, D. (1985). Introduction to Artificial Intelligence. Addison Wesley.

Choi, F. Y. (2000). Advances in domain independent linear text segmentation. In NAACL 2000, pp. 26–33.

Choi, F. Y. Y., Wiemer-Hastings, P., and Moore, J. (2001). Latent semantic analysis for text segmentation. In EMNLP 2001, pp. 109–117.

Chomsky, N. (1981). Lectures on Government and Binding. Foris, Dordrecht.

Clark, H. H. and Sengal, C. J. (1979). In search of referents for nouns and pronouns. Memory and Cognition, 7, 35–41.

Connolly, D., Burger, J. D., and Day, D. S. (1994). A machine learning approach to anaphoric reference. In Proceedings of the International Conference on New Methods in Language Processing (NeMLaP).

Corston-Oliver, S. H. (1998). Identifying the linguistic correlates of rhetorical relations. In Workshop on Discourse Relations and Discourse Markers, pp. 8–14.

Crawley, R. A., Stevenson, R. J., and Kleinman, D. (1990). The use of heuristic strategies in the interpretation of pronouns. Journal of Psycholinguistic Research, 19, 245–264.

Di Eugenio, B. (1990). Centering theory and the Italian pronominal system. In COLING-90, Helsinki, pp. 270–275.

Di Eugenio, B. (1996). The discourse functions of Italian subjects: A centering approach. In COLING-96, Copenhagen, pp. 352–357.

Filippova, K. and Strube, M. (2006). Using linguistically motivated features for paragraph boundary identification. In EMNLP 2006.

Ge, N., Hale, J., and Charniak, E. (1998). A statistical approach to anaphora resolution. In Proceedings of the Sixth Workshop on Very Large Corpora, pp. 161–171.

Gordon, P. C., Grosz, B. J., and Gilliom, L. A. (1993). Pro-

nouns, names, and the centering of attention in discourse.

Cognitive Science, 17(3), 311–347.

Grosz, B. J. (1977). The representation and use of focus in a system for understanding dialogs. In IJCAI-77, pp. 67–76. Morgan Kaufmann. Reprinted in Grosz et al. (1986).

原书第 828 页

Grosz, B. J., Joshi, A. K., and Weinstein, S. (1983). Providing a unified account of definite noun phrases in English. In ACL-83, pp. 44–50.

Grosz, B. J., Joshi, A. K., and Weinstein, S. (1995a). Centering: A framework for modeling the local coherence of discourse. Computational Linguistics, 21(2), 203–225.

Grosz, B. J., Joshi, A. K., and Weinstein, S. (1995b). Centering: A framework for modelling the local coherence of discourse. Computational Linguistics, 21(2).

Gundel, J. K., Hedberg, N., and Zacharski, R. (1993). Cognitive status and the form of referring expressions in discourse. Language, 69(2), 274–307.

Halliday, M. A. K. and Hasan, R. (1976). Cohesion in English. Longman, London. English Language Series, Title No. 9.

Haviland, S. E. and Clark, H. H. (1974). What's new? Acquiring new information as a process in comprehension. Journal of Verbal Learning and Verbal Behaviour, 13, 512–521.

Hearst, M. A. (1994). Multi-paragraph segmentation of expository text. In Proceedings of the 32nd ACL, pp. 9–16.

Hearst, M. A. (1997). Texttiling: Segmenting text into multi-paragraph subtopic passages. Computational Linguistics, 23, 33–64.

Hirschberg, J. and Litman, D. J. (1993). Empirical Studies on the Disambiguation of Cue Phrases. Computational Linguistics, 19(3), 501–530.

Hobbs, J. R. (1977). 38 examples of elusive antecedents from published texts. Tech. rep. 77-2, Department of Computer Science, City University of New York.

Hobbs, J. R. (1978). Resolving pronoun references. Lingua, 44, 311–338. Reprinted in Grosz et al. (1986).

Hobbs, J. R. (1979). Coherence and coreference. Cognitive Science, 3, 67–90.

Hobbs, J. R. (1990). Literature and Cognition. CSLI Lecture Notes 21.

Hobbs, J. R., Stickel, M. E., Appelt, D. E., and Martin, P. (1993). Interpretation as abduction. Artificial Intelligence, 63, 69–142.

Hovy, E. H. (1990). Parsimonious and profligate approaches to the question of discourse structure relations. In Proceedings of the Fifth International Workshop on Natural Language Generation, Dawson, PA, pp. 128–136.

Huls, C., Bos, E., and Classen, W. (1995). Automatic referent resolution of deictic and anaphoric expressions. Computational Linguistics, 21(1), 59–79.

Joshi, A. K. and Kuhn, S. (1979). Centered logic: The role of entity centered sentence representation in natural language inferencing. In IJCAI-79, pp. 435–439.

Joshi, A. K. and Weinstein, S. (1981). Control of inference: Role of some aspects of discourse structure – centering. In IJCAI-81, pp. 385–387.

Kameyama, M. (1986). A property-sharing constraint in centering. In ACL-86, New York, pp. 200–206.

Kan, M. Y., Klavans, J. L., and McKeown, K. R. (1998). Linear segmentation and segment significance. In Proc. 6th Workshop on Very Large Corpora (WVLC-98), Montreal, Canada, pp. 197–205.

Karamanis, N. (2003). Entity Coherence for Descriptive Text Structuring. Ph.D. thesis, University of Edinburgh.

Karamanis, N. (2007). Supplementing entity coherence with local rhetorical relations for information ordering. Journal of Logic, Language and Information. To appear.

Kawahara, T., Hasegawa, M., Shitaoka, K., Kitade, T., and Nanjio, H. (2004). Automatic indexing of lecture presentations using unsupervised learning of presumed discourse markers. Speech and Audio Processing, IEEE Transactions on, 12(4), 409–419.

Kehler, A. (1993). The effect of establishing coherence in ellipsis and anaphora resolution. In Proceedings of the 31st ACL, Columbus, Ohio, pp. 62–69.

Kehler, A. (1994a). Common topics and coherent situations: Interpreting ellipsis in the context of discourse inference. In Proceedings of the 32nd ACL, Las Cruces, New Mexico, pp. 50–57.

Kehler, A. (1994b). Temporal relations: Reference or discourse coherence?. In Proceedings of the 32nd ACL, Las Cruces, New Mexico, pp. 319–321.

Kehler, A. (1997a). Current theories of centering for pronoun interpretation: A critical evaluation. Computational Linguistics, 23(3), 467–475.

Kehler, A. (1997b). Probabilistic coreference in information extraction. In EMNLP 1997, Providence, RI, pp. 163–173.

Kehler, A. (2000). Coherence, Reference, and the Theory of Grammar. CSLI Publications.

Kehler, A., Appelt, D. E., Taylor, L., and Simma, A. (2004). The (non)utility of predicate-argument frequencies for pronoun interpretation. In HLT-NAACL-04.

Kennedy, C. and Boguraev, B. (1996). Anaphora for everyone: Pronominal anaphora resolution without a parser. In COLING-96, Copenhagen, pp. 113–118.

Knott, A. and Dale, R. (1994). Using linguistic phenomena to motivate a set of coherence relations. Discourse Processes, 18(1), 35–62.

Kozima, H. (1993). Text segmentation based on similarity between words. In Proceedings of the 31st ACL, pp. 286–288.

Lambrecht, K. (1994). Information Structure and Sentence Form. Cambridge University Press.

Lappin, S. and Leass, H. (1994). An algorithm for pronominal anaphora resolution. Computational Linguistics, 20(4), 535–561.

Lascarides, A. and Asher, N. (1993). Temporal interpretation, discourse relations, and common sense entailment. Linguistics and Philosophy, 16(5), 437–493.

Longacre, R. E. (1983). The Grammar of Discourse. Plenum Press.

原书第 829 页

Mann, W. C. and Thompson, S. A. (1987). Rhetorical structure theory: A theory of text organization. Tech. rep. RS-87-190, Information Sciences Institute.

Manning, C. D. (1998). Rethinking text segmentation models: An information extraction case study. Tech. rep. SULTRY-98-07-01, University of Sydney.

Marcu, D. (2000a). The rhetorical parsing of unrestricted texts: A surface-based approach. Computational Linguistics, 26(3), 395–448.

Marcu, D. (Ed.). (2000b). The Theory and Practice of Discourse Parsing and Summarization. MIT Press.

Marcu, D. and Echihabi, A. (2002). An unsupervised approach to recognizing discourse relations. In ACL-02, pp. 368–375.

Matthews, A. and Chodorow, M. S. (1988). Pronoun resolution in two-clause sentences: Effects of ambiguity, antecedent location, and depth of embedding. Journal of Memory and Language, 27, 245–260.

McCarthy, J. F. and Lehnert, W. G. (1995). Using decision trees for coreference resolution. In IJCAI-95, Montreal, Canada, pp. 1050–1055.

Miltsakaki, E., Prasad, R., Joshi, A. K., and Webber, B. L. (2004a). Annotating discourse connectives and their arguments. In Proceedings of the NAACL/HLT Workshop: Frontiers in Corpus Annotation.

Miltsakaki, E., Prasad, R., Joshi, A. K., and Webber, B. L. (2004b). The Penn Discourse Treebank. In LREC-04.

Mitkov, R. (2002). Anaphora Resolution. Longman.

Mitkov, R. and Boguraev, B. (Eds.). (1997). Proceedings of the ACL-97 Workshop on Operational Factors in Practical, Robust Anaphora Resolution for Unrestricted Texts, Madrid, Spain.

Morris, J. and Hirst, G. (1991). Lexical cohesion computed by the saural relations as an indicator of the structure of text. Computational Linguistics, 17(1), 21–48.

Ng, V. (2004). Learning noun phrase anaphoricity to improve coreference resolution: Issues in representation and optimization. In ACL-04.

Ng, V. (2005). Machine learning for coreference resolution: From local classification to global ranking. In ACL-05.

Ng, V. and Cardie, C. (2002a). Identifying anaphoric and nonanaphoric noun phrases to improve coreference resolution. In COLING-02.

Ng, V. and Cardie, C. (2002b). Improving machine learning approaches to coreference resolution. In ACL-02.

Nissim, M., Dingare, S., Carletta, J., and Steedman, M. (2004). An annotation scheme for information status in dialogue. In LREC-04, Lisbon.

Passonneau, R. and Litman, D. J. (1993). Intention-based segmentation: Human reliability and correlation with linguistic cues. In Proceedings of the 31st ACL, Columbus, Ohio, pp. 148–155.

Peirce, C. S. (1955). Abduction and induction. In Buchler, J. (Ed.), Philosophical Writings of Peirce, pp. 150–156. Dover Books, New York.

Pevzner, L. and Hearst, M. A. (2002). A critique and improvement of an evaluation metric for text segmentation. Computational Linguistics, 28(1), 19–36.

Poesio, M., Stevenson, R., Di Eugenio, B., and Hitzeman, J. (2004). Centering: A parametric theory and its instantiations. Computational Linguistics, 30(3), 309–363.

Poesio, M. and Vieira, R. (1998). A corpus-based investigation of definite description use. Computational Linguistics, 24(2), 183–216.

Polanyi, L., Culy, C., van den Berg, M., Thione, G. L., and Ahn, D. (2004a). A Rule Based Approach to Discourse Parsing. In Proceedings of SIGDIAL.

Polanyi, L., Culy, C., van den Berg, M., Thione, G. L., and Ahn, D. (2004b). Sentential Structure and Discourse Parsing. In Discourse Annotation Workshop, ACL04.

Polanyi, L. (1988). A formal model of the structure of discourse. Journal of Pragmatics, 12.

Power, R., Scott, D., and Bouayad-Agha, N. (2003). Document structure. Computational Linguistics, 29(2), 211–260.

Prince, E. (1981). Toward a taxonomy of given-new information. In Cole, P. (Ed.), Radical Pragmatics, pp. 223–255. Academic Press.

Prince, E. (1992). The ZPG letter: Subjects, definiteness, and information-status. In Thompson, S. and Mann, W. (Eds.), Discourse Description: Diverse Analyses of a Fundraising Text, pp. 295–325. John Benjamins, Philadelphia/Amsterdam.

Prüst, H. (1992). On Discourse Structuring, VP Anaphora, and Gapping. Ph.D. thesis, University of Amsterdam.

Reynar, J. C. (1994). An automatic method of finding topic boundaries. In Proceedings of the 32nd ACL, pp. 27–30.

Reynar, J. C. (1999). Statistical models for topic segmentation. In ACL/EACL-97, pp. 357–364.

Sanders, T. J. M., Spooren, W. P. M., and Noordman, L. G. M. (1992). Toward a taxonomy of coherence relations. Discourse Processes, 15, 1–35.

Scha, R. and Polanyi, L. (1988). An augmented context free grammar for discourse. In COLING-88, Budapest, pp. 573–577.

Sidner, C. L. (1979). Towards a computational theory of definite anaphora comprehension in English discourse. Tech. rep. 537, MIT Artificial Intelligence Laboratory, Cambridge, MA.

Sidner, C. L. (1983). Focusing in the comprehension of definite anaphora. In Brady, M. and Berwick, R. C. (Eds.), Computational Models of Discourse, pp. 267–330. MIT Press.

Smyth, R. (1994). Grammatical determinants of ambiguous pronoun resolution. Journal of Psycholinguistic Research, 23, 197–229.

Soon, W. M., Ng, H. T., and Lim, D. C. Y. (2001). A machine learning approach to coreference resolution of noun phrases. Computational Linguistics, 27(4), 521–544.

原书第 830 页

Sporleder, C. and Lapata, M. (2004). Automatic paragraph identification: A study across languages and domains. In EMNLP 2004.

Sporleder, C. and Lapata, M. (2006). Automatic Paragraph Identification: A Study across Languages and Domains. ACM Transactions on Speech and Language Processing (TSLP), 3(2).

Sporleder, C. and Lascarides, A. (2005). Exploiting Linguistic Cues to Classify Rhetorical Relations. In Proceedings of the Recent Advances in Natural Language Processing (RANLP-05), Borovets, Bulgaria.

Strube, M. and Hahn, U. (1996). Functional centering. In ACL-96, Santa Cruz, CA, pp. 270–277.

Sundheim, B. (1995). Overview of results of the MUC-6 evaluation. In Proceedings of the Sixth Message Understanding Conference (MUC-6), Columbia, MD, pp. 13–31.

van Deemter, K. and Kibble, R. (2000). On coreferring: coreference in muc and related annotation schemes. Computational Linguistics, 26(4), 629–637.

Vieira, R. and Poesio, M. (2000). An empirically based system for processing definite descriptions. Computational Linguistics, 26(4), 539–593.

Vilain, M., Burger, J. D., Aberdeen, J., Connolly, D., and Hirschman, L. (1995). A model-theoretic coreference scoring scheme. In MUC6 '95: Proceedings of the 6th Conference on Message Understanding. ACL.

Walker, M. A., Iida, M., and Cote, S. (1994). Japanese discourse and the process of centering. Computational Linguistics, 20(2).

Walker, M. A., Joshi, A. K., and Prince, E. (Eds.). (1998). Centering in Discourse. Oxford University Press.

Webber, B. L. (2004). D-LTAG: extending lexicalized TAG to discourse. Cognitive Science, 28(5), 751–79.

Webber, B. L., Knott, A., Stone, M., and Joshi, A. (1999). Discourse relations: A structural and presuppositional account using lexicalised TAG. In ACL-99, College Park, MD, pp. 41–48.

Webber, B. L. (1978). A Formal Approach to Discourse Anaphora. Ph.D. thesis, Harvard University.

Webber, B. L. (1983). So what can we talk about now?. In Brady, M. and Berwick, R. C. (Eds.), Computational Models of Discourse, pp. 331–371. The MIT Press. Reprinted in Grosz et al. (1986).

Webber, B. L. (1991). Structure and ostension in the interpretation of discourse deixis. Language and Cognitive Processes, 6(2), 107–135.

Winograd, T. (1972). Understanding Natural Language. Academic Press.

Wolf, F. and Gibson, E. (2005). Representing discourse coherence: A corpus-based analysis. Computational Linguistics, 31(2), 249–287.

Woods, W. A. (1978). Semantics and quantification in natural language question answering. In Yovits, M. (Ed.), Advances in Computers, Vol. 17, pp. 2–87. Academic Press.

Woods, W. A., Kaplan, R. M., and Nash-Webber, B. L. (1972). The Lunar Sciences Natural Language Information System: Final report. Tech. rep. 2378, Bolt, Beranek, and Newman, Inc., Cambridge, MA.

原书第 831 页

22 INFORMATION EXTRACTION

I am the very model of a modern Major-General,

I’ve information vegetable, animal, and mineral,

I know the kings of England, and I quote the fights historical

From Marathon to Waterloo, in order categorical...

Gilbert and Sullivan, Pirates of Penzance

Imagine that you are an analyst with an investment firm that tracks airline stocks. You're given the task of determining the relationship (if any) between airline announcements of fare hikes and the behavior of their stocks on the following day. Historical data about stock prices is easy to come by, but what about the information about airline announcements? To do a reasonable job on this task, you would need to know at least the name of the airline, the nature of the proposed fare hike, the dates of the announcement and possibly the response of other airlines. Fortunately, this information resides in archives of news articles reporting on airline's actions, as in the following recent example.

Citing high fuel prices, United Airlines said Friday it has increased fares by $6 per round trip on flights to some cities also served by lower-cost carriers. American Airlines, a unit of AMR Corp., immediately matched the move, spokesman Tim Wagner said. United, a unit of UAL Corp., said the increase took effect Thursday and applies to most routes where it competes against discount carriers, such as Chicago to Dallas and Denver to San Francisco.

Of course, distilling information like names, dates and amounts from naturally occurring text is a non-trivial task. This chapter presents a series of techniques that can be used to extract limited kinds of semantic content from text. This process of information extraction (IE) turns the unstructured information embedded in texts into structured data. More concretely, information extraction is an effective way to populate the contents of a relational database. Once the information is encoded formally, we can apply all the capabilities provided by database systems,

原书第 832 页

statistical analysis packages and other forms of decision support systems to address the problems we're trying to solve.

As we proceed through this chapter, we'll see that robust solutions to IE problems are actually clever combinations of techniques we've seen earlier in the book. In particular, the finite-state methods described in Chs. 2 and 3, the probabilistic models introduced in Chs. 4 through 6 and the syntactic chunking methods from Ch. 13 form the core of most current approaches to information extraction. Before diving into the details of how these techniques are applied, let's quickly introduce the major problems in IE and how they can be approached.

The first step in most IE tasks is to detect and classify all the proper names mentioned in a text — a task generally referred to as named entity recognition (NER). Not surprisingly, what constitutes a proper name and the particular scheme used to classify them is application-specific. Generic NER systems tend to focus on finding the names of people, places and organizations that are mentioned in ordinary news texts; practical applications have also been built to detect everything from the names of genes and proteins (Settles, 2005) to the names of college courses (McCallum, 2005).

Our introductory example contains 13 instances of proper names, which we'll refer to as $ \underline{\text{named entity}} $ mentions, which can be classified as either organizations, people, places, times or amounts.

Having located all of the mentions of named entities in a text, it is useful to link, or cluster, these mentions into sets that correspond to the entities behind the mentions. This is the task of reference resolution, which we introduced in Ch. 21, and is also an important component in IE. In our sample text, we would like to know that the United Airlines mention in the first sentence and the United mention in the third sentence refer to the same real world entity. This general reference resolution problem also includes anaphora resolution as a sub-problem. In this case, determining that the two uses of it refer to United Airlines and United respectively.

The task of relation detection and classification is to find and classify semantic relations among the entities discovered in a given text. In most practical settings, the focus of relation detection is on small fixed sets of binary relations. Generic relations that appear in standard system evaluations include family, employment, part-whole, membership, and geospatial relations. The relation detection and classification task is the one that most closely corresponds to the problem of populating a relational database. Relation detection among entities is also closely related to the problem of discovering semantic relations among words introduced in Ch. 20.

Our sample text contains 3 explicit mentions of generic relations: United is a part of UAL, American Airlines is a part of AMR and Tim Wagner is an employee

原书第 833 页

of American Airlines. Domain-specific relations from the airline industry would include the fact that United serves Chicago, Dallas, Denver and San Francisco.

In addition to knowing about the entities in a text and their relation to one another, we might like to find and classify the events in which the entities are participating; this is the problem of event detection and classification. In our sample text, the key events are the fare increase by United and the ensuing increase by American. In addition, there are several events reporting these main events as indicated by the two uses of said and the use of cite. As with entity recognition, event detection brings with it the problem of reference resolution; we need to figure out which of the many event mentions in a text refer to the same event. In our running example, the events referred to as the move and the increase in the second and third sentences are the same as the increase in the first sentence.

The problem of figuring out when the events in a text happened and how they relate to each other in time raises the twin problems of temporal expression detection and temporal analysis. Temporal expression detection tells us that our sample text contains the temporal expressions Friday and Thursday. Temporal expressions include date expressions such as days of the week, months, holidays, etc., as well as relative expressions including phrases like two days from now or next year. They also include expressions for clock times such as noon or 3:30 PM.

The overall problem of temporal analysis is to map temporal expressions onto specific calendar dates or times of day and then to use those times to situate events in time. It includes the following subtasks.

  • Fixing the temporal expressions with respect to an anchoring date or time, typically the dateline of the story in the case of news stories;
  • Associating temporal expressions with the events in the text:

• Arranging the events into a complete and coherent timeline.

In our sample text, the temporal expressions Friday and Thursday should be anchored with respect to the dateline associated with the article itself. We also know that Friday refers to the time of United's announcement, and Thursday refers to the time that the fare increase went into effect (i.e. the Thursday immediately preceding the Friday). Finally, we can use this information to produce a timeline where United's announcement follows the fare increase and American's announcement follows both of those events. Temporal analysis of this kind is useful in nearly any NLP application that deals with meaning, including question answering, summarization and dialogue systems.

Finally, many texts describe stereotypical situations that recur with some frequency in the domain of interest. The task of template-filling is to find documents that evoke such situations and then fill the slots in templates with appropriate material. These slot-fillers may consist of text segments extracted directly from the

原书第 834 页

text, or they may consist of concepts that have been inferred from text elements via some additional processing (times, amounts, entities from an ontology, etc.).

Our airline text is an example of this kind of stereotypical situation since airlines are often attempting to raise fares and then waiting to see if competitors follow along. In this situation, we can identify United as a lead airline that initially raised its fares, $6 as the amount by which fares are being raised, Thursday as the effective date for the fare increase, and American as an airline that followed along. A filled template from our original airline story might look like the following.

FARE-RAISE ATTEMPT:

AMOUNT: $6

EFFECTIVE DATE: 2006-10-26

FOLLOWER: AMERICAN AIRLINES

The following sections will review current approaches to each of these problems in the context of generic news text. Sec. 22.5 then describes how many of these problems arise in the context of procecessing biology texts.

22.1 NAMED ENTITY RECOGNITION

NAMED ENTITY

The starting point for most information extraction applications is the detection and classification of the named entities in a text. By named entity, we simply mean anything that can be referred to with a proper name. This process of named entity recognition refers to the combined task of finding spans of text that constitute proper names and then classifying the entities being referred to according to their type.

Generic news-oriented NER systems focus on the detection of things like people, places, and organizations. Figures 22.1 and 22.2 provide lists of typical named entity types with examples of each. Specialized applications may be concerned with many other types of entities, including commercial products, weapons, works of art, or as we'll see in Sec. 22.5, proteins, genes and other biological entities. What these applications all share is a concern with proper names, the characteristic ways that such names are signaled in a given language or genre, and a fixed set of categories of entities from a domain of interest.

By the way that names are signaled, we simply mean that names are denoted in a way that sets them apart from ordinary text. For example, if we're dealing with standard English text, then two adjacent capitalized words in the middle of a text are likely to constitute a name. Further, if they are are preceded by a Dr. or followed by an MD, then it is likely that we're dealing with a person. In contrast, if they are preceded by arrived in or followed by NY then we're probably dealing

原书第 835 页
Section 22.1. Named Entity Recognition
TypeTagSample Categories
People OrganizationPERIndividuals, fictional characters, small groups
Organization LocationORGCompanies, agencies, political parties, religious groups, sports teams
Geo-Political EntityLOCPhysical extents, mountains, lakes, seas
Facility VehiclesGPECountries, states, provinces, counties
FACBridges, buildings, airports
VEHPlanes, trains and automobiles
Figure 22.1 A list of generic named entity types with the kinds of entities they refer to.
TypeExample
People OrganizationTuring is often considered to be the father of modern computer science.\nThe IPCC said it is likely that future tropical cyclones will become more intense.\nThe Mt. Sanitas loop hike begins at the base of Sunshine Canyon.\nPalo Alto is looking at raising the fees for parking in the University Avenue district
Location Geo-Political EntityFacility Drivers were advised to consider either the Tappan Zee Bridge or the Lincoln Tunnel.\nThe updated Mini Cooper retains its charm and agility.
Vehicles
Figure 22.2 Named entity types with examples.

with a location. Note that these signals include facts about the proper names as well as their surrounding contexts.

TEMPORAL EXPRESSIONS

NUMERICAL EXPRESSIONS

The notion of a named entity is commonly extended to include things that aren't entities per se, but nevertheless have practical importance and do have characteristic signatures that signal their presence; examples include dates, times, named events and other kinds of temporal expressions, as well as measurements, counts, prices and other kinds of numerical expressions. We'll consider some of these later in Sec. 22.3.

Let's revisit the sample text introduced earlier with the named entities marked (with TIME and MONEY used to mark the temporal and monetary expressions).

Citing high fuel prices, [ORG United Airlines] said [TIME Friday] it has increased fares by [MONEY $6] per round trip on flights to some cities also served by lower-cost carriers. [ORG American Airlines], a unit of [ORG AMR Corp.], immediately matched the move, spokesman [PERS Tim Wagner] said. [ORG United], a unit of [ORG UAL Corp.], said the increase took effect [TIME Thursday] and applies to most routes where it competes against discount carriers, such as [LOC Chicago] to [LOC Dallas] and [LOC Denver] to [LOC San Francisco].

原书第 836 页
Chapter 22. Information Extraction
NamePossible Categories
Washington Downing St. IRA Louis VuittonPerson, Location, Political Entity, Organization, Facility Location, OrganizationPerson, Organization, Monetary InstrumentPerson, Organization, Commercial Product
Figure 22.3 names.Common categorical ambiguities associated with various proper
[PERS Washington] was born into slavery on the farm of James Burroughs.\n[ORG Washington] went up 2 games to 1 in the four-game series.\nBlair arrived in [LOC Washington] for what may well be his last state visit.\nIn June, [GPE Washington] passed a primary seatbelt law.\nThe [FAC Washington] had proved to be a leaky ship, every passage I made...
Figure 22.4 Examples of type ambiguities in the use of the name Washington.

As shown, this text contains 13 mentions of named entities including 5 organizations, 4 locations, 2 times, 1 person, and 1 mention of money. The 5 organizational mentions correspond to 4 unique organizations, since United and United Airlines are distinct mentions that refer to the same entity.

← 21.5 Consider the following passage, from Brennan et al. (1987):22.1.1 Ambiguity in Named Entity Recognition →