← 学习库 Speech and Language Processing 本册目录

22.5.3 Biological Roles and Relations

Finding and normalizing all the mentions of biological entities in a text is a preliminary step to determining the roles played by entities in the text. Two ways to do this that have been the focus of recent research are to discover and classify the expressed binary relations between the entities in a text, and to identify and classify the roles played by entities with respect to the central events in the text. These two tasks correspond roughly to the tasks of classifying the relationship between pairs of entities as described in Sec. 22.2, and to the semantic role labeling task introduced in Ch. 20.

Consider the following example texts that express binary relations between entities.

(22.25) These results suggest that con A-induced $ [_{DISEASE} $ hepatitis] $ was ameliorated by pretreatment with $ [_{TREATMENT} $ TJ-135].

(22.26) [DISEASE Malignant mesodermal mixed tumor of the uterus] following [TREATMENT irradiation]

Each of these examples asserts a relationship between a disease and a treatment. In the first example, the relationship can be classified as that of curing. In the second example, the disease is a result of the mentioned treatment. Rosario and Hearst (2004) present a system for the classification of 7 kinds disease-treatment relations. In this work, a series of HMM-based generative models as well as a discriminative neural network model were successfully applied.

More generally, a wide-range of rule-based and statistical approaches have been applied to binary relation recognition problems such as this. Examples of other widely studied biomedical relation recognition problems include genes and their biological functions (Blaschke et al., 2005), genes and drugs (Rindflesch et al., 2000), genes and mutations (Rebholz-Schuhmann et al., 2004), and protein-protein interactions (Rosario and Hearst, 2005).

Now consider the following example that corresponds to a semantic role labeling style of problem.

[THEME Full-length cPLA2] was [TARGET phosphorylated] stoichiometrically by [AGENT p42 mitogen-activated protein (MAP) kinase] in vitro... and the major site of phosphorylation was identified by amino acid sequencing as [SITE Ser505]

The phosphorylation event that lies at the core of this text has three semantic roles associated with it: the causal AGENT of the event, the THEME or entity being phosphorylated and finally the location, or SITE of the event. The problem is to identify the constituents in the input that play these roles and assign them the correct role labels. Note that this example, contains a further complication in that the sec

原书第 874 页

and event mention phosphorylation must be identified as coreferring with the first phosphorylated in order to capture the SITE role correctly.

Much of the difficulty with semantic role labeling in the biomedical domain stems from the preponderance of nominalizations in these texts. Nominalizations like phosphorylation typically offer fewer syntactic cues to signal their arguments than their verbal equivalents, making the identification task more difficult. A further complication is that different semantic roles arguments often occur as parts of the same, or dominating nominal constituents. To see this consider the following examples.

(22.28) Serum stimulation of fibroblasts in floating matrices does not result in [TARGET [ARG1 ERK] translocation] to the [ARG3 nucleus] and there was decreased serum activation of upstream members of the ERK signaling pathway, MEK and Raf,

(22.29) The translocation of RelA/p65 was investigated using Western blotting and immunocytochemistry. the COX-2 inhibitor SC236 worked directly through suppressing $ [TARGET[ARG3\ nuclear]translocation] $ of $ [ARG1\ RelA/p65] $.

(22.30) Following UV treatment, Mcl-1 protein synthesis is blocked, the existing pool of Mcl-1 protein is rapidly degraded by the proteasome, and [ARG1 [ARG2 cytosolic] Bcl-xL] [TARGET translocates] to the [ARG3 mitochondria]

Each these examples contains arguments that are bundled into constituents with other arguments or with the target predicate itself. For example, in the second example the constituent nuclear translocation signals both the TARGET and the ARG3 role.

Both rule-based and statistical approaches have been applied to these semantic role-like problems. As with relation-finding and NER, the choice of algorithm is less important than the choice of features, many of which are derived from accurate syntactic analyses. However, since there are no large treebanks available for biological texts, we are left with the option using off-the-shelf parsers trained on generic newswire texts. Of course, the errors introduced in this process may negate whatever power we can derive from syntactic features. Therefore, an important area of research revolves around the adaptation of generic syntactic tools to this domain (Blitzer et al., 2006).

Relational and event extraction applications in this domain often have an extremely limited foci. The motivation for this is that even systems with narrow scope can make a contribution to the productivity of working bioscientists. An extreme example of this is the RLIMS-P system discussed earlier. It tackles only the verb phosphorylate and the associated nominalization phosphorylation. Nevertheless, this system was successfully used to produce a large online database that is in widespread use by the research community.

原书第 875 页

As the targets of biomedical information extraction applications have become more ambitious, the range of BioNLP application types has become correspondingly more broad. Computational lexical semantics and semantic role labelling (Verspoor et al., 2003; Wattarujeekrit et al., 2004; Ogren et al., 2004; Kogan et al., 2005; Cohen and Hunter, 2006), summarization (Lu et al., 2006), and question-answering are all active research topics in the biomedical domain. Shared tasks like BioCreative continue to be a source of large data sets for named entity recognition, question-answering, relation extraction, and document classification (Hirschman and Blaschke, 2006), as well as a venue for head-to-head assessment of the benefits of various approaches to information extraction tasks.

22.6 SUMMARY

This chapter has explored a series of techniques for extracting limited forms of semantic content from texts. Most techniques can be characterized as problems in detection followed by classification.

  • Named entities can be recognized and classified by statistical sequence labeling techniques.
  • Relations among entities can be detected and classified using supervised learning methods when annotated training data is available; lightly supervised bootstrapping methods can be used when small numbers of seed tuples or seed patterns are available.
  • Reasoning about time can be facilitated by detecting and normalizing temporal expressions through a combination of statistical learning and rule-based methods.
  • Rule-based and statistical methods can be used to detect, classify and order events in time. The TimeBank corpus can facilitate the training and evaluation of temporal analysis systems.
  • Template-filling applications can recognize stereotypical situations in texts and assign elements from the text to roles represented as fixed sets of slots.
  • Information extraction techniques have proven to be particularly effective in processing texts from the biological domain.

• Scripts, plans and goals...

原书第 876 页

BIBLIOGRAPHICAL AND HISTORICAL NOTES

The earliest work on information extraction addressed the template-filling task and was performed in the context of the Frump system (DeJong, 1982). Later work was stimulated by the U.S. government sponsored MUC conferences (Sundheim, 1991, 1992, 1993, 1995). Chinchor et al. (1993) describes the evaluation techniques used in the MUC-3 and MUC-4 conferences. Hobbs (1997) partially credits the inspiration for FASTUS to the success of the University of Massachusetts CIRCUS system (Lehnert et al., 1991) in MUC-3. The SCISOR system is another system based loosely on cascades and semantic expectations that did well in MUC-3 (Jacobs and Rau, 1990).

Due to the difficulty of reusing or porting systems from one domain to another, attention shifted to the problem of automatic knowledge acquisition for these systems. The earliest supervised learning approaches to IE are described in Cardie (1993), Cardie (1994), Riloff (1993), Soderland et al. (1995), Huffman (1996), and Freitag (1998).

These early learning efforts focused on automating the knowledge acquisition process for mostly finite-state rule-based systems. Their success, and the earlier success of HMM-based methods for automatic speech recognition, led to the development of statistical systems based on sequence labeling. Early efforts applying HMMs to IE problems include the work of Bikel et al. (1997, 1999) and Freitag and McCallum (1999). Subsequent efforts demonstrated the effectiveness of a range of statistical methods including MEMMs (McCallum et al., 2000), CRFs (Lafferty et al., 2001) and SVMs (Sassano and Utsuro, 2000; McNamee and Mayfield, 2002).

Progress in this area continues to be stimulated by formal evaluations with shared benchmark datasets. The MUC evaluations of the mid-1990s were succeeded by the Automatic Content Extraction (ACE) program evaluations held periodically from 2000 to 2007. $ ^{6} $ These evaluations focused on the named entity recognition, relation detection, and temporal expression detection and normalization tasks. Other IE evaluations include the 2002 and 2003 CoNLL shared tasks on language-independent named entity recognition (Sang, 2002; Sang and De Meulder, 2003), and the 2007 SemEval tasks on temporal analysis (Verhagen et al., 2007) and people search (Artiles et al., 2007).

The scope of information extraction continues to expand to meet the ever-increasing needs of applications for novel kinds of information. Some of the emerging IE tasks that we haven't discussed include the classification of gender

原书第 877 页

(Koppel et al., 2002), moods (Mishne and de Rijke, 2006), sentiment, affect and opinions (Qu et al., 2004). Much of this work involves user-generated content in the context of social media such as blogs, discussion forums, newsgroups and the like. Research results in this domain have been the focus of a number of recent workshops and conferences (Nicolov et al., 2006; Nicolov and Glance, 2007).

EXERCISES

22.1 Develop a set of regular expressions to recognize the character shape features described in Fig. 22.7.

22.2 Using a statistical sequence modeling toolkit of your choosing, develop and evaluate an NER system.

22.3 The IOB labeling scheme given in this chapter isn't the only possible one. For example, an E tag might be added to mark the end of entities, or the B tag can be reserved only for those situations where an ambiguity exists between adjacent entities. Propose a new set of IOB tags for use with your NER system. Perform experiments and compare its performance against the scheme presented in this chapter.

22.4 Names of works of art (books, movies, video games, etc.) are quite different from the kinds of named entities we've discussed in this chapter. Collect a list of names of works of art from a particular category from a web-based source (eg. gutenberg.org, amazon.com, imdb.com, etc.). Analyze your list and give examples of ways that the names in it are likely to be problematic for the techniques described in this chapter.

22.5 Develop an NER system specific to the category of names that you collected in the last exercise. Evaluate your system on a collection of text likely to contain instances of these named entities.

22.6 Acronym expansion, the process of associating a phrase with a particular acronym, can be accomplished by a simple form of relational analysis. Develop a system based on the relation analysis approaches described in this chapter to populate a database of acronym expansions. If you focus on English Three Letter Acronyms (TLAs) you can evaluate your system's performance by comparing it to Wikipedia's TLA page (en.wikipedia.org/wiki/Category:Lists_of_TLAs).

22.7 Collect a corpus of biographical Wikipedia entries of prominent people from some coherent area of interest (sports, business, computer science, linguistics, etc.).

原书第 878 页

Develop a system that can extract an occupational timeline for the subjects of these articles. For example, the Wikipedia entry for Peter Norvig might result in the ordering: Sun, Harlequin, Junglee, NASA, Google; the entry for David Beckham would be: Manchester United, Real Madrid, Los Angeles Galaxy.

22.8 A useful functionality in newer email and calendar applications is the ability to associate temporal expressions associated with events in emails (doctor's appointments, meeting planning, party invitations, etc.) with specific calendar entries. Collect a corpus of emails containing temporal expressions related to event planning. How do these expressions compare to the kind of expressions commonly found in news text that we've been discussing in this chapter?

22.9 Develop and evaluate a recognition system capable of recognizing temporal expressions of the kind appearing in your email corpus.

22.10 Design a system capable of normalizing these expressions to the degree required to insert them into a standard calendaring application.

22.11 Acquire the CMU seminar announcement corpus and develop a template-filling system using any of the techniques mentioned in Sec. 22.4. Analyze how well your system performs as compared to state-of-the-art results on this corpus.

22.12 Develop a new template that covers a situation commonly reported on by standard news sources. Carefully characterize your slots in terms of the kinds of entities that appear as slot-fillers. Your first step in this exercise should be to acquire a reasonably sized corpus of stories that instantiate your template.

22.13 Given your corpus, develop an approach to annotating the relevant slots in your corpus so that it can serve as a training corpus. Your approach should involve some hand-annotation, but should not be based solely on it.

22.14 Retrain your system and analyze how well it functions on your new domain.

22.15 Species identification is a critical issue for biomedical information extraction applications such as document routing and classification. But it is especially crucial for realistic versions of the gene normalization problem.

Build a species identification system that works on the document level, using the machine learning or rule-based method of your choice. As gold standard data, use the BioCreative gene normalization data (biocreative.sourceforge.net).

22.16 Build, or borrow, a named entity recognition system that targets mentions of genes and gene products in texts. As development data, use the BioCreative gene mention corpus (biocreative.sourceforge.net).

原书第 879 页

22.17 Build a gene normalization system that maps the output of your gene mention recognition system to the appropriate database entry. Use the BioCreative gene normalization data as your development and test data, be sure you don't give your system access to the species identification in the metadata.

原书第 880 页

Agichtein, E. and Gravano, L. (2000). Snowball: Extracting relations from large plain-text collections. In Proceedings of the 5th ACM International Conference on Digital Libraries.

Ahn, D., Adafre, S. F., and de Rijke, M. (2005). Extracting temporal information from open domain text: A comparative exploration. In Proceedings of the 5th Dutch-Belgian Information Retrieval Workshop (DIR'05).

Allen, J. (1984). Towards a general theory of action and time. Artificial Intelligence, 23(2), 123–154.

Appelt, D. E., Hobbs, J. R., Bear, J., Israel, D., Kameyama, M., Kehler, A., Martin, D., Myers, K., and Tyson, M. (1995). SRI International FASTUS system MUC-6 test results and analysis. In Proceedings of the Sixth Message Understanding Conference (MUC-6), San Francisco, pp. 237–248. Morgan Kaufmann.

Appelt, D. E. and Israel, D. (1997). ANLP-97 tutorial: Building information extraction systems. Available as www.ai.sri.com/~appelt/ie-tutorial/.

Artiles, J., Gonzalo, J., and Sekine, S. (2007). The semeval-2007 weps evaluation: Establishing a benchmark for the web people search task. In Proceedings of the Fourth International Workshop on Semantic Evaluations (SemEval-2007), Prague, Czech Republic, pp. 64–69. Association for Computational Linguistics.

Ashburner, M., Ball, C. A., Blake, J. A., Botstein, D., Butler, H., Cherry, J. M., Davis, A. P., Dolinski, K., Dwight, S. S., Eppig, J. T., Harris, M. A., Hill, D. P., Issel-Tarver, L., Kasarskis, A., Lewis, S., Matese, J. C., Richardson, J. E., Ringwald, M., Rubin, G. M., and Sherlock, G. (2000). Gene ontology: tool for the unification of biology. Nature Genetics, 25(1), 25–29.

Bikel, D. M., Miller, S., Schwartz, R., and Weischedel, R. (1997). Nymble: a high-performance learning namefinder. In Proceedings of ANLP-97, pp. 194–201.

Bikel, D. M., Schwartz, R., and Weischedel, R. (1999). An algorithm that learns what's in a name. Machine Learning, 34, 211–231.

Blaschke, C., Leon, E. A., Krallinger, M., and Valencia, A. (2005). Evaluation of BioCreative assessment of task 2. BMC Bioinformatics, 6(2).

Blitzer, J., McDonald, R., and Pereira, F. (2006). Domain adaptation with structural correspondence learning. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, Sydney, Australia.

Bunescu, R. C. and Mooney, R. J. (2005). A shortest path dependency kernel for relation extraction. In Proceedings of the conference on Human Language Technology and Empirical Methods in Natural Language Processing, pp. 724–731.

Cardie, C. (1993). A case-based approach to knowledge acquisition for domain specific sentence analysis. In AAAI-93, pp. 798–803. AAAI Press.

Cardie, C. (1994). Domain-Specific Knowledge Acquisition for Conceptual Sentence Analysis. Ph.D. thesis, University of Massachusetts, Amherst, MA. Available as CMPSCI Technical Report 94-74.

Chinchor, N., Hirschman, L., and Lewis, D. L. (1993). Evaluating Message Understanding systems: An analysis of the third Message Understanding Conference. Computational Linguistics, 19(3), 409–449.

Cohen, K. B. and Hunter, L. (2006). A critical review of PASBio's argument structures for biomedical verbs. BMC Bioinformatics, 7(Suppl 3).

Cohen, K. B., Dolbey, A., Mensah, A. G., and Hunter, L. (2002). Contrast and variability in gene names. In Proceedings of the ACL Workshop on Natural Language Processing in the Biomedical Domain, pp. 14–20.

Cohen, K. B. and Hunter, L. (2004). Natural language processing and systems biology. In Dubitzky, W. and Azuaje, F. (Eds.), Artificial Intelligence Methods and Tools for Systems Biology, pp. 147–174. Springer.

Culotta, A. and Sorensen, J. (2004). Dependency tree kernels for relation extraction. In Proceedings of the 42nd Annual Meeting on Association for Computational Linguistics.

DeJong, G. F. (1982). An overview of the FRUMP system. In Lehnert, W. G. and Ringle, M. H. (Eds.), Strategies for Natural Language Processing, pp. 149–176. Lawrence Erlbaum.

Etzioni, O., Cafarella, M., Downey, D., Popescu, A.-M., Shaked, T., Soderland, S., Weld, D. S., and Yates, A. (2005). Unsupervised named-entity extraction from the web: An experimental study. Artificial Intelligence, 165(1), 91–134.

Ferro, L., Gerber, L., Mani, I., Sundheim, B., and Wilson, G. (2005). Tides 2005 standard for the annotation of temporal expressions. Tech. rep., MITRE.

Fisher, D., Soderland, S., McCarthy, J., Feng, F., and Lehnert, W. G. (1995). Description of the UMass system as used for MUC-6. In Proceedings of the Sixth Message

原书第 881 页

Understanding Conference (MUC-6), San Francisco, pp. 127–140. Morgan Kaufmann.

Freitag, D. (1998). Multistrategy learning for information extraction. In ICML 1998, Madison, WI, pp. 161–169.

Freitag, D. and McCallum, A. (1999). Information extraction using hmms and shrinkage. In Proceedings of the AAAI-99 Workshop on Machine Learning for Information Retrieval.

Gaizauskas, R., Wakao, T., Humphreys, K., Cunningham, H., and Wilks, Y. (1995). University of Sheffield: Description of the LaSIE system as used for MUC-6. In Proceedings of the Sixth Message Understanding Conference (MUC-6), San Francisco, pp. 207–220. Morgan Kaufmann.

Grishman, R. and Sundheim, B. (1995). Design of the MUC-6 evaluation. In Proceedings of the Sixth Message Understanding Conference (MUC-6), San Francisco, pp. 1–11. Morgan Kaufmann.

Hirschman, L. and Blaschke, C. (2006). Evaluation of text mining in biology. In Ananiadou, S. and McNaught, J. (Eds.), Text Mining for Biology and Biomedicine, chap. 9, pp. 213–245. Artech House, Norwood, MA.

Hobbs, J. R., Appelt, D. E., Bear, J., Israel, D., Kameyama, M., Stickel, M. E., and Tyson, M. (1997). FASTUS: A cascaded finite-state transducer for extracting information from natural-language text. In Roche, E. and Schabes, Y. (Eds.), Finite-State Language Processing, pp. 383–406. MIT Press.

Huffman, S. (1996). Learning information extraction patterns from examples. In Wertmer, S., Riloff, E., and Scheller, G. (Eds.), Connectionist, Statistical, and Symbolic Approaches to Learning Natural Language Processing, pp. 246–260. Springer, Berlin.

ISO8601 (2004). Data elements and interchange formats information interchange representation of dates and times. Tech. rep., International Organization for Standards (ISO).

Jackson, P. and Moulinier, I. (2002). Natural language processing for online applications: text retrieval, extraction, and categorization. John Benjamins Publishing Company.

Jacobs, P. and Rau, L. (1990). SCISOR: A system for extracting information from on-line news. Communications of the ACM, 33(11), 88–97.

Jr., W. A. B., Cohen, K. B., Fox, L., Acquaah-Mensah, G. K., and Hunter, L. (2007). Manual curation is not

sufficient for annotation of genomic databases. Bioinformatics, 23, i41–i48.

Jr., W. A. B., Lu, Z., Johnson, H. L., Caporaso, J. G., Paquette, J., Lindemann, A., White, E. K., Medvedeva, O., Cohen, K. B., and Hunter, L. (2006). An integrated approach to concept recognition in biomedical text. In Proceedings of BioCreative 2006.

Kinoshita, S., Cohen, K. B., Ogren, P. V., and Hunter, L. (2005). BioCreAtIvE Task1A: entity identification with a stochastic tagger. BMC Bioinformatics, 6(1).

Kogan, Y., Collier, N., Pakhomov, S., and Krauthammer, M. (2005). Towards semantic role labeling & IE in the medical literature. In AMIA 2005 Symposium Proceedings, pp. 410–414.

Koppel, M., Argamon, S., and Shimoni, A. R. (2002). Automatically categorizing written texts by author gender. Literary and Linguistic Computing, 17(4), 401–412.

Lafferty, J. D., McCallum, A., and Pereira, F. C. N. (2001). Conditional random fields: Probabilistic models for segmenting and labeling sequence data. In ICML 2001, Stanford, CA.

Lehnert, W. G., Cardie, C., Fisher, D., Riloff, E., and Williams, R. (1991). Description of the CIRCUS system as used for MUC-3. In Sundheim, B. (Ed.), Proceedings of the Third Message Understanding Conference, pp. 223–233. Morgan Kaufmann.

Lu, Z., Cohen, B. K., and Hunter, L. (2006). Finding GeneRIFs via Gene Ontology annotations.. In PSB 2006, pp. 52–63.

McCallum, A. (2005). Information extraction: Distilling structured data from unstructured text. ACM Queue, 48–57.

McCallum, A., Freitag, D., and Pereira, F. C. N. (2000). Maximum entropy Markov models for information extraction and segmentation. In ICML 2000, pp. 591–598.

McNamee, P. and Mayfield, J. (2002). Entity extraction without language-specific resources. In Proceedings of the Conference on Computational Natural Language Learning (CoNLL-2002), Taipei, Taiwan.

Mikheev, A., Moens, M., and Grover, C. (1999). Named entity recognition without gazetteers. In Proceedings of the Ninth Conference of the European Chapter of the Association for Computational Linguistics, Morristown, NJ, USA, pp. 1–8. Association for Computational Linguistics.

原书第 882 页

Mishne, G. and de Rijke, M. (2006). MoodViews: Tools for blog mood analysis. In Nicolov, N., Salvetti, F., Liberman, M., and Martin, J. H. (Eds.), Computational Approaches to Analyzing Weblogs: Papers from the 2006 Spring Symposium, Stanford, Ca. AAAI.

Morgan, A. A., Wellner, B., Colombe, J. B., Arens, R., Colosimo, M. E., and Hirschman, L. (2007). Evaluating human gene and protein mention normalization to unique identifiers. In Pacific Symposium on Biocomputing, pp. 281–291.

Ng, S.-K. (2006). Integrating text mining with data mining. In Ananiadou, S. and McNaught, J. (Eds.), Text mining for biology and biomedicine. Artech House Publishers.

Nicolov, N. and Glance, N. (Eds.). (2007). Proceedings of the First International Conference on Weblogs and Social Media (ICWSM), Boulder, CO.

Nicolov, N., Salvetti, F., Liberman, M., and Martin, J. H. (Eds.). (2006). Computational Approaches to Analyzing Weblogs: Papers from the 2006 Spring Symposium, Stanford, Ca. AAAI.

Ogren, P. V., Cohen, K. B., Acquaah-Mensah, G. K., Eberlein, J., and Hunter, L. (2004). The compositional structure of Gene Ontology terms. In Pac Symp Biocomput, pp. 214–225.

Peshkin, L. and Pfefer, A. (2003). Bayesian information extraction network. In Proceedings of the Eighteenth International Joint Conference on Artificial Intelligence.

Pustejovsky, J., Castao, J., Inqria, R., Saur, R., Gaizauskas, R., Setzer, A., and Katz, G. (2003a). TimeML: robust specification of event and temporal expressions in text. In Proceedings of the Fifth International Workshop on Computational Semantics (IWCS-5).

Pustejovsky, J., Hanks, P., Saur, R., See, A., Gaizauskas, R., Setzer, A., Radev, D., Sundheim, B., Day, D., Ferro, L., and Lazo, M. (2003b). The TIMEBANK corpus. In Proceedings of Corpus Linguistics 2003 Conference, pp. 647–656.

Pustejovsky, J., Ingría, R., Sauri, R., Castano, J., Littman, J., Gaizauskas, R., Setzer, A., Katz, G., and Mani, I. (2005). The Specification Language TimeML, chap. 27. Oxford, Oxford, England.

Qu, Y., Shanahan, J., and Wiebe, J. (Eds.). (2004). Exploring Attitude and Affect in Text: Papers from the 2004 Spring Symposium, Stanford, Ca. AAAI.

Rebholz-Schuhmann, D., Marcel, S., Albert, S., Tolle, R., Casari, G., and Kirsch, H. (2004). Automatic extraction of mutations from medline and cross-validation with omim. Nucleic Acids Research, 32(1), 135–142.

Riloff, E. (1993). Automatically constructing a dictionary for information extraction tasks. In AAAI-93, Washington, D.C., pp. 811–816.

Riloff, E. and Jones, R. (1999). Learning dictionaries for information extraction by multi-level bootstrapping. In Proceedings of the Sixteenth National Conference on Artificial Intelligence (AAAI), pp. 474–479.

Rindflesch, T. C., Tanabe, L., Weinstein, J. N., and Hunter, L. (2000). EDGAR: Extraction of drugs, genes and relations from the biomedical literature. In Pacific Symposium on Biocomputing, pp. 515–524.

Rosario, B. and Hearst, M. A. (2004). Classifying semantic relations in bioscience texts. In Proceedings of ACL 2004, pp. 430–437.

Rosario, B. and Hearst, M. A. (2005). Multi-way Relation Classification: Application to Protein-Protein Interactions. In Proceedings of the 2005 HLT-NAACL.

Roth, D. and tau Yih, W. (2001). Relational learning via propositional algorithms: An information extraction case study. In Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI), pp. 1257–1263.

Sang, E. F. T. K. (2002). Introduction to the conll-2002 shared task: Language-independent named entity recognition. In Proceedings of CoNLL-2002, pp. 155–158. Taipei, Taiwan.

Sang, E. F. T. K. and De Meulder, F. (2003). Introduction to the conll-2003 shared task: Language-independent named entity recognition. In Daelemans, W. and Osborne, M. (Eds.), Proceedings of CoNLL-2003, pp. 142–147. Edmonton, Canada.

Sassano, M. and Utsuro, T. (2000). Named entity chunking techniques in supervised learning for Japanese named entity recognition. In COLING-00, Saarbrcken, Germany, pp. 705–711.

Schank, R. C. and Abelson, R. P. (1977). Scripts, Plans, Goals and Understanding. Lawrence Erlbaum.

Schwartz, A. S. and Hearst, M. A. (2003). A simple algorithm for identifying abbreviation definitions in biomedical text. In Pacific Symposium on Biocomputing, Vol. 8, pp. 451–462.

Settles, B. (2005). ABNER: an open source tool for automatically tagging genes, proteins and other entity names in text. Bioinformatics, 21(14), 3191–3192.

原书第 883 页

Soderland, S., Fisher, D., Aseltine, J., and Lehnert, W. G. (1995). CRYSTAL: Inducing a conceptual dictionary. In IJCAI-95, Montreal, pp. 1134–1142.

Sundheim, B. (Ed.). (1991). Proceedings of the Third Message Understanding Conference. Morgan Kaufmann.

Sundheim, B. (Ed.). (1992). Proceedings of the Fourth Message Understanding Conference. Morgan Kaufmann.

Sundheim, B. (Ed.). (1993). Proceedings, Fifth Message Understanding Conference (MUC-5), Baltimore, MD. Morgan Kaufmann.

Sundheim, B. (Ed.). (1995). Proceedings of the Sixth Message Understanding Conference. Morgan Kaufmann.

van Rijsbergen, C. J. (1975). Information Retrieval. Butterworths, London.

Verhagen, M., Gaizauskas, R., Schilder, F., Hepple, M., Katz, G., and Pustejovsky, J. (2007). Semeval-2007 task 15: Tempeval temporal relation identification. In Proceedings of the Fourth International Workshop on Semantic Evaluations (SemEval-2007), Prague, Czech Republic, pp. 75–80. Association for Computational Linguistics.

Verspoor, C. M., Joslyn, C., and Papcun, G. J. (2003). The gene ontology as a source of lexical semantic knowledge for a biological natural language processing application. In Proceedings of the SIGIR'03 Workshop on Text Analysis and Search for Bioinformatics, Toronto, CA.

Wattarujeekrit, T., Shah, P. K., and Collier, N. (2004). PASBio: predicate-argument structures for event extraction in molecular biology. BMC Bioinformatics, 5(155).

Weischedel, R. (1995). BBN: Description of the PLUM system as used for MUC-6. In Proceedings of the Sixth Message Understanding Conference (MUC-6), San Francisco, pp. 55–70. Morgan Kaufmann.

Yeh, A., Morgan, A., Colosimo, M., and Hirschman, L. (2005). BioCreative task 1A: gene mention finding evaluation. BMC Bioinformatics, 6(1).

Zhou, G., Su, J., Zhang, J., and Zhang, M. (2005). Exploring various knowledge in relation extraction. In Proceedings of the 43rd Annual Meeting of the Association for Computational Linguistics (ACL'05), Ann Arbor, Michigan, pp. 427–434. Association for Computational Linguistics.

原书第 884 页

23 QUESTION ANSWERING AND SUMMARIZATION

‘Alright’, said Deep Thought. ‘The Answer to the Great Question...’ ‘Yes!’

'Of Life The Universe and Everything...' said Deep Thought.

'Yes!'

'Is...'

'Yes...!!!!...?'

'Forty-two', said Deep Thought, with infinite majesty and calm... Douglas Adams, The Hitchhiker's Guide to the Galaxy

I read War and Peace... It's about Russia... Woody Allen, Without Feathers

Because so much text information is available generally on the Web, or in specialized collections such as PubMed, or even on the hard drives of our laptops, the single most important use of language processing these days is to help us query and extract meaning from these large repositories. If we have a very structured idea of what we are looking for, we can use the information extraction algorithms of the previous chapter. But many times we have an information need that is best expressed more informally in words or sentences, and we want to find either a specific answer fact, or a specific document, or something in between.

In this chapter we introduce the tasks of question answering (QA) and summarization, tasks which produce specific phrases, sentences, or short passages, often in response to a user's need for information expressed in a natural language query. In studying these topics, we will also cover highlights from the field of information retrieval (IR), the task of returning documents which are relevant to a particular natural language query. IR is a complete field in its own right, and we will only be giving a brief introduction to it here, but one that is essential for understand QA and summarization.

In this chapter we focus on a central idea behind all of these subfields, the idea of meeting a user's information needs by extracting passages directly from documents or from document collections like the Web.

Information retrieval (IR) is an extremely broad field, encompassing a wide-range of topics pertaining to the storage, analysis, and retrieval of all manner of media,

原书第 885 页

including text, photographs, audio, and video (Baeza-Yates and Ribeiro-Neto, 1999). Our concern in this chapter is solely with the storage and retrieval of text documents in response to users' word-based queries for information. In section 23.1 we present the vector space model, some variant of which is used in most current systems, including most web search engines.

Rather than make the user read through an entire document, we'd often prefer to give a single concise short answer. Researchers have been trying to automate this process of question answering since the earliest days of computational linguistics (Simmons, 1965).

(23.1) Who founded Virgin Airlines?

(23.2) What is the average age of the onset of autism?

The simplest form of question answering is dealing with factoid questions. As the name implies, the answers to factoid questions are simple facts that can be found in short text strings. The following are canonical examples of this kind of question.

(23.3) Where is Apple Computer based?

Each of these questions can be answered directly with a text string that contain the name of person, a temporal expression, or a location, respectively. Factoid questions, therefore, are questions whose answers can be found in short spans of text and correspond to a specific, easily characterized, category, often a named entity of the kind we discussed in Ch. 22. These answers may be found on the Web, or alternatively within some smaller text collection. For example a system might answer questions about a company's product line by searching for answers in documents on a particular corporate website or internal set of documents. Effective techniques for answering these kinds of questions are described in Sec. 23.2.

Sometimes we are seeking information whose scope is greater than a single factoid, but less than an entire document. In such cases we might need a summary of a document or set of documents. The goal of text summarization is to produce an abridged version of a text which contains the important or relevant information. For example we might want to generate an abstract of a scientific article, a summary of email threads, a headline for a news article, or generate the short snippets that web search engines like Google return to the user to describe each retrieved document. For example, Fig. 23.1 shows some sample snippets from Google summarizing the first four documents returned from the query German Expressionism Brücke.

To produce these various kinds of summaries, we'll introduce algorithms for summarizing single documents, and those for producing summaries of multiple documents by combining information from different textual sources.

Finally, we turn to a field that tries to go beyond factoid question answering by borrowing techniques from summarization to try to answer more complex questions like the following:

(23.4) Who is Celia Cruz?

(23.5) What is a Hajj?

(23.6) In children with an acute febrile illness, what is the efficacy of single-medication therapy with acetaminophen or ibuprofen in reducing fever?

原书第 886 页
Image
Figure 23.1 The first 4 snippets from Google for German Expressionism Brücke.

Answers to questions such as these do not consist of simple named entity strings. Rather they involve potentially lengthy coherent texts that knit together an array of associated facts to produce a biography, a complete definition, a summary of current events, or a comparison of clinic results on particular medical interventions. In addition to the complexity and style differences in these answers, the facts that go into such answers may be context, user, and time dependent.

Current methods answer these kinds of complex questions by piecing together relevant text segments that come from summarizing longer documents. For example, we might construct an answer from text segments extracted from a corporate report, a set of medical research journal articles, or a set of relevant news articles or web pages. This idea of summarizing text in response to a user query is called query-based summarization or focused summarization, and will be explored in Sec. 23.5.

Finally, we reserve for Ch. 24 all discussion of the role that questions play in extended dialogues; this chapter focuses only on responding to a single query.

23.1 INFORMATION RETRIEVAL

INFORMATION RETRIEVAL

Information retrieval (IR) is a growing field that encompasses a wide range of topics related to the storage and retrieval of all manner of media. The focus of this section is with the storage of text documents and their subsequent retrieval in response to users' requests for information. In this section our goal is just to give a sufficient overview of information retrieval techniques to lay a foundation for the following sections on question answering and summarization. Readers with more interest specifically in information retrieval should see the references at the end of the chapter.

Most current information retrieval systems are based on a kind of extreme version of compositional semantics in which the meaning of a document resides solely in the

原书第 887 页

set of words it contains. To revisit the Mad Hatter's quote from the beginning of Ch. 19, in these systems I see what I eat and I eat what I see mean precisely the same thing. The ordering and constituency of the words that make up the sentences that make up documents play no role in determining their meaning. Because they ignore syntactic information, these approaches are often referred to as bag-of-words models.

Before moving on, we need to introduce some new terminology. In information retrieval, a document refers generically to the unit of text indexed in the system and available for retrieval. Depending on the application, a document can refer to anything from intuitive notions like newspaper articles, or encyclopedia entries, to smaller units such as paragraphs and sentences. In web-based applications, it can refer to a web page, a part of a page, or to an entire website. A collection refers to a set of documents being used to satisfy user requests. A term refers to a lexical item that occurs in a collection, but it may also include phrases. Finally, a query represents a user's information need expressed as a set of terms.

Image

The specific information retrieval task that we will consider in detail is known as ad hoc retrieval. In this task, it is assumed that an unaided user poses a query to a retrieval system, which then returns a possibly ordered set of potentially useful documents. The high level architecture is shown in Fig. 23.2.

Figure 23.2 The architecture of an ad hoc IR system.
← 22.5.2 Gene Normalization23.1.1 The Vector Space Model →