18.5.4 Prepositional Phrases
At a fairly abstract level, prepositional phrases serve two distinct functions: they assert binary relations between their heads and the constituents to which they are attached, and they signal arguments to constituents that have an argument structure. These two functions argue for two distinct types of prepositional phrases that differ based on their semantic attachments. We will consider three places in the grammar where prepositional phrases serve these roles: modifiers of noun phrases, modifiers of verb phrases, and arguments to verb phrases.
Nominal Modifier Prepositional Phrases
Modifier prepositional phrases denote a binary relation between the concept being modified, which is external to the prepositional phrase, and the head of the prepositional phrase. Consider the following example and its associated meaning representation:
(18.20) A restaurant on Broadway.
$$ \exists x\,Isa(x,Restaurant)\land On(x,Pearl) $$
The relevant grammar rules that govern this example are the following:
NP $ \rightarrow $ Det Nominal
Nominal $ \rightarrow $ Nominal PP
$$ PP\to PN P $$
Proceeding in a bottom-up fashion, the semantic attachment for this kind of relational preposition should provide a two-place predicate with its arguments distributed over two $ \lambda $-expressions, as in the following:
$$ P\to on\qquad\{\lambda y\lambda x\,On(x,y)\} $$
With this kind of arrangement, the first argument to the predicate is provided by the head of prepositional phrase and the second is provided by the constituent that the prepositional phrase is ultimately attached to. The following semantic attachment provides the first part:
$$ PP\to P N P\quad\{P.sem(N P.sem)\} $$
This $ \lambda $-application results in a new $ \lambda $-expression where the remaining argument is the inner $ \lambda $-variable.
This remaining argument can be incorporated using the following nominal construction:
Nominal $\rightarrow$ Nominal PP $\left\{\lambda z\mathrm{Nominal}.\mathrm{sem}(z)\land\mathrm{PP}.\mathrm{sem}(z)\right\}$
Verb Phrase Modifier Prepositional Phrases
The general approach to modifying verb phrases is similar to that of modifying nominals. The differences lie in the details of the modification in the verb phrase rule; the attachments for the preposition and prepositional phrase rules are unchanged. Let's consider the phrase ate dinner in a hurry which is governed by the following verb phrase rule:
$$ VP\to VP PP $$
The meaning representation of the verb phrase constituent in this construction, ate dinner, is a $ \lambda $-expression where the $ \lambda $-variable represents the as yet unseen subject.
$$ \lambda x\exists e\,Isa(e,Eating)\land Eater(e,x)\land Eaten(e,Dinner) $$
The representation of the prepositional phrase is also a $ \lambda $-expression where the $ \lambda $-variable is the second argument in the PP semantics.
$$ \lambda x\;In(x,<\exists h\;Hurry(h)>) $$
The correct representation for the modified verb phrase should contain the conjunction of these two representations with the Eating event variable filling the first argument slot of the In expression. In addition, this modified representation must remain a $ \lambda $-expression with the unbound Eater variable as the new $ \lambda $-variable. The following attachment expression fulfills all of these requirements:
$$ VP\rightarrow VP\;PP\qquad\left\{\lambda y VP.sem(y)\land PP.sem(VP.sem.variable)\right\} $$
There are two aspects of this attachment that require some elaboration. The first involves the application of the constituent verb phrases’ $ \lambda $-expression to the variable y. Binding the lower $ \lambda $-expression’s variable to a new variable allows us to lift the lower variable to the level of the newly created $ \lambda $-expression. The result of this technique is a new $ \lambda $-expression with a variable that, in effect, plays the same role as the original variable in the lower expression. In this case, this allows a $ \lambda $-expression to be modified during the analysis process before the argument to the expression is actually available.
The second notable aspect of this attachment involves the VP.sem.variable notation. This notation is used to access the event-variable representing the underlying meaning of the verb phrase, in this case, e. This is analogous to the notation used to provide access to the various parts of complex-terms introduced earlier.
Applying this attachment to the current example yields the following representation, which is suitable for combination with a subsequent subject noun phrase:
$$ \begin{array}{c}\lambda y\exists e Isa(e,Eating)\land Eater(e,y)\land Eaten(e,Dinner)\\\land In(e,<\exists h Hurry(h)>)\end{array} $$
Verb Argument Prepositional Phrases
The prepositional phrases in this category serve to signal the role an argument plays in some larger event structure. As such, the preposition itself does not actually modify the meaning of the noun phrase. Consider the following example of role signaling prepositional phrases:
(18.21) I need to go from Boston to Dallas.
In examples like this, the arguments of go are expressed as prepositional phrases. However, the meaning representations of these phrases should consist solely of the unaltered representation of their head nouns. To handle this, argument prepositional phrases are treated in the same way that non-branching grammatical rules are; the semantic attachment of the noun phrase is copied unchanged to the semantics of the larger phrase.
$$ PP\to PN\quad\{NP.sem\} $$
The verb phrase can then assign this meaning representation to the appropriate event role. A more complete account of how these argument bearing prepositional phrases map to underlying event roles will be presented in Ch. 19.
18.6 INTEGRATING SEMANTIC ANALYSIS INTO THE EARLEY PARSER
In Section 18.1, we suggested a simple pipeline architecture for a semantic analyzer where the results of a complete syntactic parse are passed to a semantic analyzer. The motivation for this notion stems from the fact that the compositional approach requires the syntactic parse before it can proceed. It is, however, also possible to perform semantic analysis in parallel with syntactic processing. This is possible because in our compositional framework, the meaning representation for a constituent can be created as soon as all of its constituent parts are present. This section describes just such an approach to integrating semantic analysis into the Earley parser from Ch. 13.
The integration of semantic analysis into an Earley parser is straightforward and follows precisely the same lines as the integration of unification into the algorithm given in Ch. 16. Three modifications are required to the original algorithm:
1. The rules of the grammar are given a new field to contain their semantic attachments.
2. The states in the chart are given a new field to hold the meaning representation of the constituent.
3. The ENQUEUE function is altered so that when a complete state is entered into the chart its semantics are computed and stored in the state's semantic field.
Figure 18.6 shows ENQUEUE modified to create meaning representations. When ENQUEUE is passed a complete state that can successfully unify its unification constraints it calls APPLY-SEMANTICS to compute and store the meaning representation for this state. Note the importance of performing feature-structure unification prior to semantic analysis. This ensures that semantic analysis will be
| procedure ENQUEUE(state, chart-entry)\nif INCOMPLETE?(state) then\n if state is not already in chart-entry then\n PUSH(state, chart-entry)\n\n else if UNIFY-STATE(state) succeeds then\n if APPLY-SEMANTICS(state) succeeds then\n if state is not already in chart-entry then\n PUSH(state, chart-entry) |
| procedure APPLY-SEMANTICS(state)\nmeaning-rep ← APPLY(state.semantic-attachment,state)\nif meaning-rep does not equal failure then\nstate.meaning-rep ← meaning-rep |
| Figure 18.6 The ENQUEUE function modified to handle semantics. If the state is complete and unification succeeds then ENQUEUE calls APPLY-SEMANTICS to compute and store the meaning representation of completed states. |
performed only on valid trees and that features needed for semantic analysis will be present.
The primary advantage of this integrated approach over the pipeline approach lies in the fact that APPLY-SEMANTICS can fail in a manner similar to the way that unification can fail. If a semantic ill-formedness is found in the meaning representation being created, the corresponding state can be blocked from entering the chart. In this way, semantic considerations can be brought to bear during syntactic processing. Ch. 19 describes in some detail the various ways that this notion of ill-formedness can be realized.
Unfortunately, this also illustrates one of the primary disadvantages of integrating semantics directly into the parser—considerable effort may be spent on the semantic analysis of orphan constituents that do not in the end contribute to a successful parse. The question of whether the gains made by bringing semantics to bear early in the process outweigh the costs involved in performing extraneous semantic processing can only be answered on a case-by-case basis.
18.7 IDIOMS AND COMPOSITIONALITY
Ce corps qui s'appélait et qui s'appelle encore le saint empire romain n'était en aucune manière ni saint, ni romain, ni empire.
This body, which called itself and still calls itself the Holy Roman Empire, was neither Holy, nor Roman, nor an Empire.
Voltaire $ ^{2} $, 1756
As innocuous as it seems, the principle of compositionality runs into trouble fairly quickly when real language is examined. There are many cases where the meaning of a constituent is not based on the meaning of its parts, at least not in the straightforward compositional sense. Consider the following WSJ examples:
(18.22) Coupons are just the tip of the iceberg.
(18.23) The SEC's allegations are only the tip of the iceberg.
(18.24) Coronary bypass surgery, hip replacement and intensive-care units are but the tip of the iceberg.
The phrase the tip of the iceberg in each of these examples clearly doesn't have much to do with tips or icebergs. Instead, it roughly means something like the beginning. The most straightforward way to handle idiomatic constructions like these is to introduce new grammar rules specifically designed to handle them. These idiomatic rules mix lexical items with grammatical constituents, and introduce semantic content that is not derived from any of its parts. Consider the following rule as an example of this approach:
NP $ \rightarrow $ the tip of the iceberg
{Beginning}
The lower case items on the right-hand side of this rule are intended to represent precisely words in the input. Although, the constant Beginning should not be taken too seriously as a meaning representation for this idiom, it does illustrate the idea that the meaning of this idiom is not based on the meaning of any of its parts. Note that an Earley-style analyzer with this rule will now produce two parses when this phrase is encountered: one representing the idiom and one representing the compositional meaning.
As with the rest of the grammar, it may take a few tries to get these rules right. Consider the following iceberg examples from the WSJ corpus:
(18.25) And that's but the tip of Mrs. Ford's iceberg.
(18.26) These comments describe only the tip of a 1,000-page iceberg.
(18.27) The 10 employees represent the merest tip of the iceberg.
The rule given above is clearly not general enough to handle these cases. These examples indicate that there is a vestigial syntactic structure to this idiom that permits some variation in the determiners used, and also permits some adjectival modification of both the iceberg and the tip. A more promising rule would be something like the following:
{Beginning}
Here the categories TipNP and IcebergNP can be given an internal nominal-like structure that permits some adjectival modification and some variation in the determiners, while still restricting the heads of these noun phrases to the lexical items tip and iceberg. Note that this syntactic solution ignores the thorny issue that the modifiers mere and 1000-page seem to indicate that both the tip and iceberg may in fact play some compositional role in the meaning of the idiom. We will return to this topic in Ch. 19, when we take up the issue of metaphor.
To summarize, handling idioms requires at least the following changes to the general compositional framework:
• Allow the mixing of lexical items with traditional grammatical constituents.
- Allow the creation of additional idiom-specific constituents needed to handle the correct range of productivity of the idiom.
- Permit semantic attachments that introduce logical terms and predicates that are not related to any of the constituents of the rule.
This discussion is obviously only the tip of an enormous iceberg. Idioms are far more frequent and far more productive than is generally recognized and pose serious difficulties for many applications, including, as we will see in Ch. 24, machine translation.
18.8 SUMMARY
This chapter explores the notion of syntax-driven semantic analysis. Among the highlights of this chapter are the following topics:
- Semantic analysis is the process whereby meaning representations are created and assigned to linguistic inputs.
- Semantic analyzers that make use of static knowledge from the lexicon and grammar can create context-independent literal, or conventional, meanings.
- The Principle of Compositionality states that the meaning of a sentence can be composed from the meanings of its parts.
- In Syntax-driven semantic analysis, the parts are the syntactic constituents of an input.
• Compositional creation of FOL formulas is possible with a few notational extensions including $ \lambda $-expressions and complex-terms.
• Compositional creation of FOL formulas is also possible using the mechanisms provided by feature structures and unification.
- Natural language quantifiers introduce a kind of ambiguity that is difficult to handle compositionally. Complex-terms can be used to compactly encode this ambiguity.
- Idiomatic language defies the principle of compositionality but can easily be handled by adapting the techniques used to design grammar rules and their semantic attachments.
BIBLIOGRAPHICAL AND HISTORICAL NOTES
As noted earlier, the principle of compositionality is traditionally attributed to Frege; Janssen (1997) discusses this attribution. Using the categorial grammar framework described in Ch. 14, Montague (1973) demonstrated that a compositional approach could be systematically applied to an interesting fragment of natural language. The rule-to-rule hypothesis was first articulated by Bach (1976). On the computational side of things, Woods's LUNAR system (Woods, 1977) was based on a pipelined syntax-first compositional analysis. Schubert and Pelletier (1982) developed an incremental rule-to-rule system based on Gazdar's GPSG approach (Gazdar, 1981, 1982; Gazdar et al., 1985). Main and Benson (1983) extended Montague's approach to the domain of question-answering.
In one of the all-too-frequent cases of parallel development, researchers in programming languages developed essentially identical compositional techniques to aid in the design of compilers. Specifically, Knuth (1968) introduced the notion of attribute grammars that associate semantic structures with syntactic structures in a one-to-one correspondence. As a consequence, the style of semantic attachments used in this chapter will be familiar to users of the YACC-style (Johnson and Lesk, 1978) compiler tools.
Semantic Grammars are due to Burton (Brown and Burton, 1975). Similar notions developed around the same time included Pragmatic Grammars (Woods, 1977) and Performance Grammars (Robinson, 1975). All centered around the notion of reshaping syntactic grammars to serve the needs of semantic processing. It
is safe to say that most modern systems developed for use in limited domains make use of some form of semantic grammar.
Most of the techniques used in the fragment of English presented in Section 18.5 are adapted from SRI's Core Language Engine (Alshawi, 1992). Additional bits and pieces were adapted from Woods (1977), Schubert and Pelletier (1982), and Gazdar et al. (1985). Of necessity, a large number of important topics were not covered in this chapter. See Alshawi (1992) for the standard gap-threading approach to semantic interpretation in the presence of long-distance dependencies. ter Meulen (1995) presents an modern treatment of tense, aspect, and the representation of temporal information. Extensive coverage of approaches to quantifier scoping can be found in Hobbs and Shieber (1987) and Alshawi (1992). van Lehn (1978) presents a set of human preferences for quantifier scoping. Over the years, a considerable amount of effort has been directed toward the interpretation of compound nominals. Linguistic research on this topic can be found in Lees (1970), Downing (1977), Levi (1978), and Ryder (1994), more computational approaches are described in Gershman (1977), Finin (1980), McDonald (1982), Pierre (1984), Arens et al. (1987), Wu (1992), Vanderwende (1994), and Lauer (1995).
There is a long and extensive literature on idioms. Fillmore et al. (1988) describe a general grammatical framework called Construction Grammar that places idioms at the center of its underlying theory. Makkai (1972) presents an extensive linguistic analysis of many English idioms. Hundreds of idiom dictionaries for second-language learners are also available. On the computational side, Becker (1975) was among the first to suggest the use of phrasal rules in parsers. Wilensky and Arens (1980) were among the first to successfully make use of this notion in their PHRAN system. Zernik (1987) demonstrated a system that could learn such phrasal idioms in context. A collection of papers on computational approaches to idioms appeared in (Fass et al., 1992).
Finally, we have skipped an entire branch of semantic analysis in which expectations driven from deep meaning representations drive the analysis process. Such systems avoid the direct representation and use of syntax, rarely making use of anything resembling a parse tree. Some of the earliest and most successful efforts along these lines were developed by Simmons (1973, 1978, 1983) and (Wilks, 1975a, 1975b). A series of similar approaches were developed by Roger Schank and his students (Riesbeck, 1975; Birnbaum and Selfridge, 1981; Riesbeck, 1986). In these approaches, the semantic analysis process is guided by detailed procedures associated with individual lexical items. The CIRCUS information extraction system (Lehnert et al., 1991) traces its roots to these systems.
EXERCISES
18.1 The attachment given on page 23 for handling noun phrases with complex determiners is not general enough to handle most possessive noun phrases. Specifically, it doesn't work for phrases like the following:
a. My sister's flight
b. My fiance's mother's flight
Create a new set of semantic attachments to handle cases like these.
18.2 Develop a set of grammar rules and semantic attachments to handle predicate adjectives such as the one following:
a. Flight 308 from New York is expensive.
b. Murphy's restaurant is cheap.
18.3 None of the attachments given in this chapter provide temporal information. Augment a small number of the most basic rules to add temporal information along the lines sketched in Ch. 17. Use your rules to create meaning representations for the following examples:
a. Flight 299 departed at 9 o'clock.
b. Flight 208 will arrive at 3 o'clock.
c. Flight 1405 will arrive late.
18.4 As noted in Ch. 17, the present tense in English can be used to refer to either the present or the future. However, it can also be used to express habitual behavior, as in the following:
Flight 208 leaves at 3 o'clock.
This could be a simple statement about today's Flight 208, or alternatively it might state that this flight leaves at 3 o'clock every day. Create a FOPC meaning representation along with appropriate semantic attachments for this habitual sense.
18.5 Implement an Earley-style semantic analyzer based on the discussion on page 30.
18.6 It has been claimed that it is not necessary to explicitly list the semantic attachment for most grammar rules. Instead, the semantic attachment for a rule should be inferable from the semantic types of the rule's constituents. For example, if a rule has two constituents, where one is a single argument $ \lambda $-expression and the
other is a constant, then the semantic attachment should obviously apply the $ \lambda $-expression to the constant. Given the attachments presented in this chapter, does this type-driven semantics seem like a reasonable idea?
18.7 Add a simple type-driven semantics mechanism to the Earley analyzer you implemented for Exercise 18.5.
18.8 Using a phrasal search on your favorite Web search engine, collect a small corpus of the tip of the iceberg examples. Be certain that you search for an appropriate range of examples (i.e., don't just search for “the tip of the iceberg”) Analyze these examples and come up with a set of grammar rules that correctly accounts for them.
18.9 Collect a similar corpus of examples for the idiom miss the boat. Analyze these examples and come up with a set of grammar rules that correctly accounts for them.
18.10 There are now a fair number of Web-based natural language question answering services that purport to provide answers to questions on a wide range of topics (see the book's Web page for pointers to current services). Develop a corpus of questions for some general domain of interest and use it to evaluate one or more of these services. Report your results. What difficulties did you encounter in applying the standard evaluation techniques to this task?
18.11 Collect a small corpus of weather reports from your local newspaper or the Web. Based on an analysis of this corpus, create a set of frames sufficient to capture the semantic content of these reports.
18.12 Implement and evaluate a small information extraction system for the weather report corpus you collected for the last exercise.
Alshawi, H. (Ed.). (1992). The Core Language Engine. MIT Press.
Arens, Y., Granacki, J., and Parker, A. (1987). Phrasal analysis of long noun sequences. In ACL-87, Stanford, CA, pp. 59–64. ACL.
Bach, E. (1976). An extension of classical transformational grammar. In Problems of Linguistic Metatheory (Proceedings of the 1976 Conference). Michigan State University.
Becker (1975). The phrasal lexicon. In Schank, R. and Nash-Webber, B. L. (Eds.), Theoretical Issues in Natural Language Processing. Cambridge, MA.
Birnbaum, L. and Selfridge, M. (1981). Conceptual analysis of natural language. In Schank, R. C. and Riesbeck, C. K. (Eds.), Inside Computer Understanding: Five Programs plus Miniatures, pp. 318–353. Lawrence Erlbaum.
Brown, J. S. and Burton, R. R. (1975). Multiple representations of knowledge for tutorial reasoning. In Bobrow, D. G. and Collins, A. (Eds.), Representation and Understanding, pp. 311–350. Academic Press.
Downing, P. (1977). On the creation and use of English compound nouns. Language, 53(4), 810–842.
Fass, D., Martin, J. H., and Hinkelman, E. A. (Eds.). (1992). Computational Intelligence: Special Issue on Non-Literal Language, Vol. 8. Blackwell, Cambridge, MA.
Fillmore, C. J., Kay, P., and O'Connor, M. C. (1988). Regularity and idiomaticity in grammatical constructions: The case of Let Alone. Language, 64(3), 510–538.
Finin, T. (1980). The semantic interpretation of nominal compounds. In AAAI-80, Stanford, CA, pp. 310–312.
Gazdar, G. (1981). Unbounded dependencies and coordinate structure. Linguistic Inquiry, 12(2), 155–184.
Gazdar, G. (1982). Phrase structure grammar. In Jacobson, P. and Pullum, G. K. (Eds.), The Nature of Syntactic Representation, pp. 131–186. Reidel, Dordrecht.
Gazdar, G., Klein, E., Pullum, G. K., and Sag, I. A. (1985). Generalized Phrase Structure Grammar. Basil Blackwell, Oxford.
Gershman, A. V. (1977). Conceptual analysis of noun groups in English. In IJCAI-77, Cambridge, MA, pp. 132–138.
Hobbs, J. R. and Shieber, S. M. (1987). An algorithm for generating quantifier scopings. Computational Linguistics, 13(1), 47–55.
Janssen, T. M. V. (1997). Compositionality. In van Benthem, J. and ter Meulen, A. (Eds.), Handbook of Logic and Language, chap. 7, pp. 417–473. North-Holland, Amsterdam.
Johnson, S. C. and Lesk, M. E. (1978). Language development tools. Bell System Technical Journal, 57(6), 2155–2175.
Knuth, D. E. (1968). Semantics of context-free languages. Mathematical Systems Theory, 2(2), 127–145.
Lauer, M. (1995). Corpus statistics meet the noun compound. In ACL-95, Cambridge, MA, pp. 47–54.
Lees, R. (1970). Problems in the grammatical analysis of English nominal compounds. In Bierwitsch, M. and Heidolph, K. E. (Eds.), Progress in Linguistics, pp. 174–187. Mouton, The Hague.
Lehnert, W. G., Cardie, C., Fisher, D., Riloff, E., and Williams, R. (1991). Description of the CIRCUS system as used for MUC-3. In Sundheim, B. (Ed.), Proceedings of the Third Message Understanding Conference, pp. 223–233. Morgan Kaufmann.
Levi, J. (1978). The Syntax and Semantics of Complex Nominals. Academic Press.
Main, M. G. and Benson, D. B. (1983). Denotational semantics for natural language question-answering programs. American Journal of Computational Linguistics, 9(1), 11–21.
Makkai, A. (1972). Idiom Structure in English. Mouton, The Hague.
McDonald, D. B. (1982). Understanding Noun Compounds. Ph.D. thesis, Carnegie Mellon University, Pittsburgh, PA. CMU Technical Report CS-82-102.
Montague, R. (1973). The proper treatment of quantification in ordinary English. In Thomason, R. (Ed.), Formal Philosophy: Selected Papers of Richard Montague, pp. 247–270. Yale University Press, New Haven, CT.
Pierre, I. (1984). Another look at nominal compounds. In COLING-84, Stanford, CA, pp. 509–516.
Riesbeck, C. K. (1975). Conceptual analysis. In Schank, R. C. (Ed.), Conceptual Information Processing, pp. 83–156. American Elsevier, New York.
Riesbeck, C. K. (1986). From conceptual analyzer to direct memory access parsing: An overview. In Advances in Cognitive Science 1, pp. 236–258. Ellis Horwood, Chichester.
Robinson, J. J. (1975). Performance grammars. In Reddy, D. R. (Ed.), Speech Recognition: Invited Paper Presented at the 1974 IEEE Symposium, pp. 401–427. Academic Press.
Ryder, M. E. (1994). Ordered Chaos: The Interpretation of English Noun-Noun Compounds. University of California Press, Berkeley.
Schubert, L. K. and Pelletier, F. J. (1982). From English to logic: Context-free computation of ‘conventional’ logical translation. American Journal of Computational Linguistics, 8(1), 27–44.
Sills, D. L. and Merton, R. K. (Eds.). (1991). Social Science Quotations. MacMillan, New York.
Simmons, R. F. (1973). Semantic networks: Their computation and use for understanding English sentences. In Schank, R. C. and Colby, K. M. (Eds.), Computer Models of Thought and Language, pp. 61–113. W.H. Freeman and Co., San Francisco.
Simmons, R. F. (1978). Rule-based computations on English. In Waterman, D. A. and Hayes-Roth, F. (Eds.), Pattern-Directed Inference Systems. Academic Press.
Simmons, R. F. (1983). Computations from the English. Prentice Hall.
ter Meulen, A. (1995). Representing Time in Natural Language. MIT Press.
van Lehn, K. (1978). Determining the scope of English quantifiers. Master's thesis, MIT, Cambridge, MA. MIT Technical Report AI-TR-483.
Vanderwende, L. (1994). Algorithm for the automatic interpretation of noun sequences. In COLING-94, Kyoto, pp. 782–788.
Wilensky, R. and Arens, Y. (1980). PHRAN: A knowledge-based natural language understander. In ACL-80, Philadelphia, PA, pp. 117–121. ACL.
Wilks, Y. (1975a). An intelligent analyzer and understander of English. Communications of the ACM, 18(5), 264–274.
Wilks, Y. (1975b). A preferential, pattern-seeking, semantics for natural language inference. Artificial Intelligence, 6(1), 53–74.
Woods, W. A. (1977). Lunar rocks in natural English: Explorations in natural language question answering. In Zampolli, A. (Ed.), Linguistic Structures Processing, pp. 521–569. North Holland, Amsterdam.
Wu, D. (1992). Automatic Inference: A Probabilistic Basis for Natural Language Interpretation. Ph.D. thesis, University of California, Berkeley, Berkeley, CA. UCB/CSD 92-692.
Zernik, U. (1987). Strategies in Language Acquisition: Learning Phrases from Examples in Context. Ph.D. thesis, University of California, Los Angeles, Computer Science Department, Los Angeles, CA.
19
LEXICAL SEMANTICS
“When I use a word”, Humpty Dumpty said in rather a scornful tone, “it means just what I choose it to mean – neither more nor less.”
How many legs does a dog have if you call its tail a leg?
Four.
Lewis Carroll, Alice in Wonderland
Calling a tail a leg doesn't make it one.
LEXEME
LEXICON
Attributed to Abraham Lincoln
The previous two chapters focused on the representation of meaning representations for entire sentences. In those discussions, we made a simplifying assumption by representing word meanings as unanalyzed symbols like EAT or JOHN or RED. But representing the meaning of a word by capitalizing it is a pretty unsatisfactory model. In this chapter we introduce a richer model of the semantics of words, drawing on the linguistic study of word meaning, a field called lexical semantics.
Before we try to define word meaning in the next section, we first need to be clear on what we mean by word, since we have used the word word in many different ways in this book.
We can use the word lexeme to mean a pairing of a particular form (orthographic or phonological) with its meaning, and a lexicon is a finite list of lexemes. For the purposes of lexical semantics, particularly for dictionaries and thesauruses, we represent a lexeme by a lemma. A lemma or citation form is the grammatical form that is used to represent a lexeme. This is often the base form; thus carpet is the lemma for carpets. The lemma or citation form for sing, sang, sung is sing. In many languages the infinitive form is used as the lemma for the verb; thus in Spanish dormir ‘to sleep’ is the lemma for verb forms like duermes ‘you sleep’. The specific forms sung or carpets or sing or duermes are called wordforms.
LEMMATIZATION
The process of mapping from a wordform to a lemma is called lemmatization. Lemmatization is not always deterministic, since it may depend on the context. For example, the wordform found can map to the lemma find (meaning 'to locate') or the lemma found ('to create an institution'), as illustrated in the following WSJ examples:
(19.1) He has looked at 14 baseball and football stadiums and found that only one-private Dodger Stadium – brought more money into a city than it took out.
(19.2) Culturally speaking, this city has increasingly displayed its determination to found the sort of institutions that attract the esteem of Eastern urbanites.
In addition, lemmas are part-of-speech specific; thus the wordform tables has two possible lemmas, the noun table and the verb table.
One way to do lemmatization is via the morphological parsing algorithms of Ch. 3. Recall that morphological parsing takes a surface form like cats and produces cat + PL. But a lemma is not necessarily the same as the stem from the morphological parse. For example, the morphological parse of the word celebrations might produce the stem celebrate with the affixes -ion and -s, while the lemma for celebrations is the longer form celebration. In general lemmas may be larger than morphological stems (e.g., New York or throw up). The intuition is that we want to have a different lemma whenever we need to have a completely different dictionary entry with its own meaning representation; we expect to have celebrations and celebration share an entry, since the difference in their meanings is mainly just grammatical, but not necessarily to share one with celebrate.
In the remainder of this chapter, when we refer to the meaning (or meanings) of a 'word', we will generally be referring to a lemma rather than a wordform.
Now that we have defined the locus of word meaning, we will proceed to different ways to represent this meaning. In the next section we introduce the idea of word sense as the part of a lexeme that represents word meaning. In following sections we then describe ways of defining and representing these senses, as well as introducing the lexical semantic aspects of the events defined in Ch. 17.
19.1 WORD SENSES
The meaning of a lemma can vary enormously given the context. Consider these two uses of the lemma bank, meaning something like ‘financial institution’ and ‘sloping mound’, respectively:
(19.3) Instead, a bank can hold the investments in a custodial account in the client's name.
(19.4) But as agriculture burgeons on the east bank, the river will shrink even more.
We represent some of this contextual variation by saying that the lemma bank has two senses. A sense (or word sense) is a discrete representation of one aspect of the meaning of a word. Loosely following lexicographic tradition, we will represent each sense by placing a superscript on the orthographic form of the lemma as in bank $ ^{1} $ and bank $ ^{2} $.
The senses of a word might not have any particular relation between them; it may be almost coincidental that they share an orthographic form. For example, the financial institution and sloping mound senses of bank seem relatively unrelated. In such cases we say that the two senses are homonyms, and the relation between the senses is one of homonymy. Thus bank $ ^{1} $ ('financial institution') and bank $ ^{2} $ ('sloping mound') are homonyms.
Sometimes, however, there is some semantic connection between the senses of a word. Consider the following WSJ 'bank' example:
While some banks furnish sperm only to married women, others are much less restrictive.
Although this is clearly not a use of the ‘sloping mound’ meaning of bank, it just as clearly is not a reference to a promotional giveaway at a financial institution. Rather, bank has a whole range of uses related to repositories for various biological entities, as in blood bank, egg bank, and sperm bank. So we could call this ‘biological repository’ sense bank $ ^{3} $. Now this new sense bank $ ^{3} $ has some sort of relation to bank $ ^{1} $; both bank $ ^{1} $ and bank $ ^{3} $ are repositories for entities that can be deposited and taken out; in bank $ ^{1} $ the entity is money, where in bank $ ^{3} $ the entity is biological.
When two senses are related semantically, we call the relationship between them polysemy rather than homonymy. In many cases of polysemy the semantic relation between the senses is systematic and structured. For example consider yet another sense of bank, exemplified in the following sentence:
The bank is on the corner of Nassau and Witherspoon.
This sense, which we can call $ \mathbf{bank}^{4} $, means something like ‘the building belonging to a financial institution’. It turns out that these two kinds of senses (an organization, and the building associated with an organization) occur together for many other words as well (school, university, hospital, etc). Thus there is a systematic relationship between senses that we might represent as
BUILDING $ \leftrightarrow $ ORGANIZATION
This particular subtype of polysemy relation is often called metonymy. Metonymy is the use of one aspect of a concept or entity to refer to other aspects of the entity, or to the entity itself. Thus we are performing metonymy when we use the phrase the White House to refer to the administration whose office is in the White House.
Other common examples of metonymy include the relation between the following pairings of senses:
• Author (Jane Austen wrote Emma) $ \leftrightarrow $ Works of Author (I really love Jane Austen)
• Animal (The chicken was domesticated in Asia) ↔ Meat (The chicken was overcooked)
Tree (Plums have beautiful blossoms) $ \leftrightarrow $ Fruit (I ate a preserved plum yesterday)
While it can be useful to distinguish polysemy from homonymy, there is no hard threshold for ‘how related’ two senses have to be to be considered polysemous. Thus the difference is really one of degree. This fact can make it very difficult to decide how many senses a word has, i.e., whether to make separate senses for closely related usages. There are various criteria for deciding that the differing uses of a word should be represented as distinct discrete senses. We might consider two senses discrete if
they have independent truth conditions, different syntactic behavior, independent sense relations, or exhibit antagonistic meanings.
Consider the following uses of the verb serve from the WSJ corpus:
(19.7) They rarely serve red meat, preferring to prepare seafood, poultry or game birds.
(19.8) He served as U.S. ambassador to Norway in 1976 and 1977.
(19.9) He might have served his time, come out and led an upstanding life.
The serve of serving red meat and that of serving time clearly have different truth conditions and presuppositions; the serve of serve as ambassador has the distinct subcategorization structure serve as NP. These heuristic suggests that these are probably three distinct senses of serve. One practical technique for determining if two senses are distinct is to conjoin two uses of a word in a single sentence; this kind of conjunction of antagonistic readings is called zeugma. Consider the following ATIS examples:
(19.10) Which of those flights serve breakfast?
(19.11) Does Midwest Express serve Philadelphia?
(19.12) Does Midwest Express serve breakfast and Philadelphia?
We use (?) to mark example those that are semantically ill-formed. The oddness of the invented third example (a case of zeugma) indicates there is no sensible way to make a single sense of serve work for both breakfast and Philadelphia. We can use this as evidence that serve has two different senses in this case.
Dictionaries tend to use many fine-grained senses so as to capture subtle meaning differences, a reasonable approach given that traditional role of dictionaries in aiding word learners. For computational purposes, we often don't need these fine distinctions and so we may want to group or cluster the senses; we have already done this for some of the examples in this chapter.
We generally reserve the word homonym for two senses which share both a pronunciation and an orthography. A special case of multiple senses that causes problems especially for speech recognition and spelling correction is homophones. Homophones are senses that are linked to lemmas with the same pronunciation but different spellings, such as wood/would or to/two/too. A related problem for speech synthesis are homographs (Ch. 8). Homographs are distinct senses linked to lemmas with the same orthographic form but different pronunciations, such as these homographs of bass:
(19.13) The expert angler from Dora, Mo., was fly-casting for bass rather than the traditional trout.
(19.14) The curtain rises to the sound of angry dogs baying and ominous bass chords sounding.
How can we define the meaning of a word sense? Can we just look in a dictionary? Consider the following fragments from the definitions of right, left, red, and blood from the American Heritage Dictionary (Morris, 1985).
right adj. located nearer the right hand esp. being on the right when facing the same direction as the observer.
left adj. located nearer to this side of the body than the right.
red n. the color of blood or a ruby.
blood n. the red liquid that circulates in the heart, arteries and veins of animals.
Note the amount of circularity in these definitions. The definition of right makes two direct references to itself, while the entry for left contains an implicit self-reference in the phrase this side of the body, which presumably means the left side. The entries for red and blood avoid this kind of direct self-reference by instead referencing each other in their definitions. Such circularity is, of course, inherent in all dictionary definitions; these examples are just extreme cases. For humans, such entries are still useful since the user of the dictionary has sufficient grasp of these other terms to make the entry in question sensible.
For computational purposes, one approach to defining a sense is to make use of a similar approach to these dictionary definitions; defining a sense via its relationship with other senses. For example, the above definitions make it clear that right and left are similar kinds of lemmas that stand in some kind of alternation, or opposition, to one another. Similarly, we can glean that red is a color, it can be applied to both blood and rubies, and that blood is a liquid. Sense relations of this sort are embodied in on-line databases like WordNet. Given a sufficiently large database of such relations, many applications are quite capable of performing sophisticated semantic tasks (even if they do not really know their right from their left).
A second computational approach to meaning representation is to create a small finite set of semantic primitives, atomic units of meaning, and then create each sense definition out of these primitives. This approach is especially common when defining aspects of the meaning of events such as semantic roles.
We will explore both of these approaches to meaning in this chapter. In the next section we introduce various relations between senses, followed by a discussion of WordNet, a sense relation resource. We then introduce a number of meaning representation approaches based on semantic primitives such as semantic roles.
19.2 RELATIONS BETWEEN SENSES
This section explores some of the relations that hold among word senses, focusing on a few that have received significant computational investigation: synonymy, antonymy, and hypernymy, as well as a brief mention of other relations like meronymy.