19.4.6 Selectional Restrictions
Semantic roles gave us a way to express some of the semantics of an argument in its relation to the predicate. In this section we turn to another way to express semantic constraints on arguments. A selectional restriction is a kind of semantic type constraint that a verb imposes on the kind of concepts that are allowed to fill its argument roles. Consider the two meanings associated with the following example:
(19.49) I want to eat someplace that's close to ICSI
There are two possible parses and semantic interpretations for this sentence. In the sensible interpretation eat is intransitive and the phrase someplace that's close to ICSI is an adjunct that gives the location of the eating event. In the nonsensical speaker-as-Godzilla interpretation, eat is transitive and the phrase someplace that's close to ICSI is the direct object and the THEME of the eating, like the NP Malaysian food in the following sentences:
(19.50) I want to eat Malaysian food.
How do we know that someplace that's close to ICSI isn't the direct object in this sentence? One useful cue is the semantic fact that the THEME of EATING events tends to be something that is edible. This restriction placed by the verb eat on the filler of its THEME argument, is called a selectional restriction. A selectional restriction is a constraint on the semantic type of some argument.
Selectional restrictions are associated with senses, not entire lexemes. We can see this in the following examples of the lexeme serve:
(19.51) Well, there was the time they served green-lipped mussels from New Zealand.
(19.52) Which airlines serve Denver?
Example (19.51) illustrates the cooking sense of serve, which ordinarily restricts its THEME to be some kind of foodstuff. Example (19.52) illustrates the provides a commercial service to sense of serve, which constrains its THEME to be some type of appropriate location. We will see in Ch. 20 that the fact that selectional restrictions are associated with senses can be used as a cue to help in word sense disambiguation.
Selectional restrictions vary widely in their specificity. Note in the following examples that the verb imagine impose strict requirements on its AGENT role (restricting it to humans and other animate entities) but places very few semantic requirements on its THEME role. A verb like diagonalize, on the other hand, places a very specific constraint on the filler of its THEME role: it has to be a matrix, while the arguments of the adjectives odorless are restricted to concepts that could possess an odor.
(19.53) In rehearsal, I often ask the musicians to imagine a tennis game.
(19.54) I cannot even imagine what this lady does all day. Radon is a naturally occurring odorless gas that can't be detected by human senses.
(19.55) To diagonalize a matrix is to find its eigenvalues.
These examples illustrate that the set of concepts we need to represent selectional restrictions (being a matrix, being able to possess an oder, etc) is quite open-ended. This distinguishes selectional restrictions from other features for representing lexical knowledge, like parts-of-speech, which are quite limited in number.
Representing Selectional Restrictions
One way to capture the semantics of selectional restrictions is to use and extend the event representation of Ch. 17. Recall that the neo-Davidsonian representation of an event consists of a single variable that stands for the event, a predicate denoting the kind of event, and variables and relations for the event roles. Ignoring the issue of the $ \lambda $-structures, and using thematic roles rather than deep event roles, the semantic contribution of a verb like eat might look like the following:
$$ \exists e,x,y\;Eating(e)\land Agent(e,x)\land Theme(e,y) $$
With this representation, all we know about y, the filler of the THEME role, is that it is associated with an Eating event via the Theme relation. To stipulate the selectional restriction that y must be something edible, we simply add a new term to that effect:
$$ \exists e,x,y\;Eating(e)\land Agent(e,x)\land Theme(e,y)\land Isa(y,EdibleThing) $$
Sense 1
hamburger, beefburger --
(a fried cake of minced beef served on a bun)
=> sandwich
=> snack food
=> dish
=> nutrient, nourishment, nutrition...
=> food, nutrient
=> substance
=> matter
=> physical entity
=> entity
When a phrase like ate a hamburger is encountered, a semantic analyzer can form the following kind of representation:
$$ \begin{array}{c}\exists e,x,y Eating(e)\land Eater(e,x)\land Theme(e,y)\land Isa(y,EdibleThing)\\\land Isa(y,Hamburger)\end{array} $$
This representation is perfectly reasonable since the membership of y in the category Hamburger is consistent with its membership in the category EdibleThing, assuming a reasonable set of facts in the knowledge base. Correspondingly, the representation for a phrase such as ate a takeoff would be ill-formed because membership in an event-like category such as Takeoff would be inconsistent with membership in the category EdibleThing.
While this approach adequately captures the semantics of selectional restrictions, there are two practical problems with its direct use. First, using FOPC to perform the simple task of enforcing selectional restrictions is overkill. There are far simpler formalisms that can do the job with far less computational cost. The second problem is that this approach presupposes a large logical knowledge-base of facts about the concepts that make up selectional restrictions. Unfortunately, although such common sense knowledge-bases are being developed, none currently have the kind of scope necessary to the task.
A more practical approach is to state selectional restrictions in terms of WordNet synsets, rather than logical concepts. Each predicate simply specifies a WordNet synset as the selectional restriction on each of its arguments. A meaning representation is well-formed if the role filler word is a hyponym (subordinate) of this synset.
For our ate a hamburger example, for example, we could set the selectional restriction on the THEME role of the verb eat to the synset {food, nutrient}, glossed as any substance that can be metabolized by an animal to give energy and build tissue: Luckily, the chain of hypernyms for hamburger shown in Fig. 19.7 reveals that hamburgers are indeed food. Again, the filler of a role need not match the restriction synset exactly, it just needs to have the synset as one of its superordinates.
We can apply this approach to the THEME roles of the verbs imagine, lift and di-
agonalize, discussed earlier. Let us restrict imagine's THEME to the synset {entity}, lift's THEME to {physical entity} and diagonalize to {matrix}. This arrangement correctly permits imagine a hamburger and lift a hamburger, while also correctly ruling out diagonalize a hamburger.
Of course WordNet is unlikely to have the exactly relevant synsets to specify selectional restrictions for all possible words of English; other taxonomies may also be used. In addition, it is possible to learn selectional restrictions automatically from corpora.
We will return to selectional restrictions in Ch. 20 where we introduce the extension to selectional preferences, where a predicate can place probabilistic preferences rather than strict deterministic constraints on its arguments.
19.5 PRIMITIVE DECOMPOSITION
Back at the beginning of the chapter, we said that one way of defining a word is to decompose its meaning into a set of primitive semantics elements or features. We saw one aspect of this method in our discussion of finite lists of thematic roles (agent, patient, instrument, etc). We turn now to a brief discussion of how this kind of model, called primitive decomposition, or componential analysis, could be applied to the meanings of all words. Wierzbicka (1992, 1996) shows that this approach dates back at least to continental philosophers like Descartes and Leibniz.
Consider trying to define words like hen, rooster, or chick. These words have something in common (they all describe chickens) and something different (their age and sex). This can be represented by using semantic features, symbols which represent some sort of primitive meaning:
hen +female, +chicken, +adult
rooster -female, +chicken, +adult
chick +chicken, -adult
A number of studies of decompositional semantics, especially in the computational literature, have focused on the meaning of verbs. Consider these examples for the verb kill:
(19.56) Jim killed his philodendron.
(19.57) Jim did something to cause his philodendron to become not alive.
There is a truth-conditional ('propositional semantics') perspective from which these two sentences have the same meaning. Assuming this equivalence, we could represent the meaning of kill as:
$$ (19.58)\quad\operatorname{K I L L}(x,y)\Leftrightarrow\operatorname{C A U S E}(x,\operatorname{B E C O M E}(\operatorname{N O T(A L I V E(y)))) $$
thus using semantic primitives like do, cause, become not, and alive.
Indeed, one such set of potential semantic primitives has been used to account for some of the verbal alternations discussed in Sec. 19.4.2 (Lakoff, 1965; Dowty, 1979). Consider the following examples.
$$ (19.59)\qquad\mathrm{J o h n~o p e n e d~t h e~d o o r.}\Rightarrow\mathrm{(C A U S E(J o h n(B E C O M E(O P E N(d o o r)))))} $$
(19.60) The door opened. $ \Rightarrow $ (BECOME(OPEN(door))
(19.61) The door is open. $ \Rightarrow $ (OPEN(door))
The decompositional approach asserts that a single state-like predicate associated with open underlies all of these examples. The differences among the meanings of these examples arises from the combination of this single predicate with the primitives CAUSE and BECOME.
While this approach to primitive decomposition can explain the similarity between states and actions, or causative and non-causative predicates, it still relies on having a very large number of predicates like open. More radical approaches choose to break down these predicates as well. One such approach to verbal predicate decomposition is Conceptual Dependency (CD), a set of ten primitive predicates, shown in Fig. 19.8.
| Primitive | Definition |
| ATRANS | The abstract transfer of possession or control from one entity to another. |
| PTRANS | The physical transfer of an object from one location to another |
| MTRANS | The transfer of mental concepts between entities or within an entity. |
| MBUILD | The creation of new information within an entity. |
| PROPEL | The application of physical force to move an object. |
| MOVE | The integral movement of a body part by an animal. |
| INGEST | The taking in of a substance by an animal. |
| EXPEL | The expulsion of something from an animal. |
| SPEAK | The action of producing a sound. |
| ATTEND | The action of focusing a sense organ. |
Below is an example sentence along with its CD representation. The verb brought is translated into the two primitives ATRANS and PTRANS to indicate the fact that the waiter both physically conveyed the check to Mary and passed control of it to her. Note that CD also associates a fixed set of thematic roles with each primitive to represent the various participants in the action.
(19.62) The waiter brought Mary the check.
$$ \begin{array}{c}\exists x,y Atrans(x)\land Actor(x,Waiter)\land Object(x,Check)\land To(x,Mary)\\\land Ptrans(y)\land Actor(y,Waiter)\land Object(y,Check)\land To(y,Mary)\end{array} $$
There are also sets of semantic primitives that cover more than just simple nouns and verbs. The following list comes from Wierzbicka (1996):
substantives: I, YOU, SOMEONE, SOMETHING, PEOPLE
mental predicates: THINK, KNOW, WANT, FEEL, SEE, HEAR speech: SAY
determiners and quantifiers: THIS, THE SAME, OTHER, ONE, TWO.
actions and events: DO, HAPPEN
evaluators: GOOD, BAD
descriptors: BIG, SMALL
time: ___ WHEN, BEFORE, AFTER
space: WHERE, UNDER, ABOVE,
partonomy and taxonomy: PART (OF), KIND (OF)
movement, existence, life: MOVE, THERE IS, LIVE
metapredicates: NOT, CAN, VERY
interclausal linkers: IF, BECAUSE, LIKE
space: FAR, NEAR, SIDE, INSIDE, HERE
time: ___ A LONG TIME, A SHORT TIME, NOW
imagination and possibility: IF... WOULD, CAN, MAYBE
Because of the difficulty of coming up with a set of primitives that can represent all possible kinds of meanings, most current computational linguistic work does not use semantic primitives. Instead, most computational work tends to use the lexical relations of Sec. 19.2 to define words.
19.6 ADVANCED CONCEPTS: METAPHOR
METAPHOR We use a metaphor when we refer to and reason about a concept or domain using words and phrases whose meanings come from a completely different domain. Metaphor is similar to metonymy, which we introduced as the use of one aspect of a concept or entity to refer to other aspects of the entity. In Sec. 19.1 we introduced metonymies like the following,
(19.63) Author (Jane Austen wrote Emma) ← Works of Author (I really love Jane Austen).
in which two senses of a polysemous word are systematically related. In metaphor, by contrast, there is a systematic relation between two completely different domains of meaning.
Metaphor is pervasive. Consider the following WSJ sentence:
(19.64) That doesn't scare Digital, which has grown to be the world's second-largest computer maker by poaching customers of IBM's mid-range machines.
The verb scare means ‘to cause fear in’, or ‘to cause to lose courage’. For this sentence to make sense, it has to be the case that corporations can experience emotions like fear or courage as people do. Of course they don’t, but we certainly speak of them and reason about them as if they do. We can therefore say that this use of scare is based on a metaphor that allows us to view a corporation as a person, which we will refer to the CORPORATION AS PERSON metaphor.
This metaphor is neither novel nor specific to this use of scare. Instead, it is a fairly conventional way to think about companies and motivates the use of resuscitate, hemorrhage and mind in the following WSJ examples:
(19.65) Fuqua Industries Inc. said Triton Group Ltd., a company it helped resuscitate, has begun acquiring Fuqua shares.
(19.66) And Ford was hemorrhaging; its losses would hit $1.54 billion in 1980.
(19.67) But if it changed its $ \underline{\text{mind}} $, however, it would do so for investment reasons, the filing said.
Each of these examples reflects an elaborated use of the basic CORPORATION AS PERSON metaphor. The first two examples extend it to use the notion of health to express a corporation's financial status, while the third example attributes a mind to a corporation to capture the notion of corporate strategy.
Metaphorical constructs such as CORPORATION AS PERSON are known as conventional metaphors. Lakoff and Johnson (1980) argue that many if not most of the metaphorical expressions that we encounter every day are motivated by a relatively small number of these simple conventional schemas.
19.7 SUMMARY
This chapter has covered a wide range of issues concerning the meanings associated with lexical items. The following are among the highlights:
- Lexical semantics is the study of the meaning of words, and the systematic meaning-related connections between words.
- A word sense is the locus of word meaning; definitions and meaning relations are defined at the level of the word sense rather than wordforms as a whole.
- Homonymy is the relation between unrelated senses that share a form, while polysemy is the relation between related senses that share a form.
• Synonymy holds between different words with the same meaning.
• Hyponymy relations hold between words that are in a class-inclusion relationship.
- Semantic fields are used to capture semantic connections among groups of lexemes drawn from a single domain.
- WordNet is a large database of lexical relations for English words.
- Semantic roles abstract away from the specifics of deep semantic roles by generalizing over similar roles across classes of verbs.
- Thematic roles are a model of semantic roles based on a single finite list of roles. Other semantic role models include per-verb semantic roles lists and proto-agent/proto-patient both of which are implemented in PropBank, and per-frame role lists, implemented in FrameNet.
- Semantic selectional restrictions allow words (particularly predicates) to post constraints on the semantic properties of their argument words.
- Primitive decomposition is another way to represent the meaning of word, in terms of finite sets of sub-lexical primitives.
BIBLIOGRAPHICAL AND HISTORICAL NOTES
GENERATIVE LEXICON QUALIA STRUCTURE
Cruse (2004) is a useful introductory linguistic text on lexical semantics. Levin and Rappaport Hovav (2005) is a research survey covering argument realization and semantic roles. Lyons (1977) is another classic reference. Collections describing computational work on lexical semantics can be found in Pustejovsky and Bergler (1992), Saint-Dizier and Viegas (1995) and Klavans (1995).
The most comprehensive collection of work concerning WordNet can be found in Fellbaum (1998). There have been many efforts to use existing dictionaries as lexical resources. One of the earliest was Amsler's (1980, 1981) use of the Merriam Webster dictionary. The machine readable version of Longman's Dictionary of Contemporary English has also been used (Boguraev and Briscoe, 1989). See Pustejovsky (1995), Pustejovsky and Boguraev (1996), Martin (1986) and Copestake and Briscoe (1995), inter alia, for computational approaches to the representation of polysemy. Pustejovsky's theory of the Generative Lexicon, and in particular his theory of the qualia structure of words, is another way of accounting for the dynamic systematic polysemy of words in context.
As we mentioned earlier, thematic roles are one of the oldest linguistic models, proposed first by the Indian grammarian Panini sometimes between the 7th and 4th centuries BCE. Their modern formulation is due to Fillmore (1968) and Gruber (1965). Fillmore's work had a large and immediate impact on work in natural language processing, as much early work in language understanding used some version of Fillmore's case roles (e.g., Simmons (1973, 1978, 1983)).
Work on selectional restrictions as a way of characterizing semantic well-formedness began with Katz and Fodor (1963). McCawley (1968) was the first to point out that selectional restrictions could not be restricted to a finite list of semantic features, but had to be drawn from a larger base of unrestricted world knowledge.
Lehrer (1974) is a classic text on semantic fields. More recent papers addressing this topic can be found in Lehrer and Kittay (1992). Baker et al. (1998) describe ongoing work on the FrameNet project.
The use of semantic primitives to define word meaning dates back to Leibniz; in linguistics, the focus on componential analysis in semantics was due to ? (?). See Nida (1975) for a comprehensive overview of work on componential analysis. Wierzbicka (1996) has long been a major advocate of the use of primitives in linguistic semantics; Wilks (1975) has made similar arguments for the computational use of primitives in machine translation and natural language understanding. Another prominent effort has been Jackendoff's Conceptual Semantics work (1983, 1990), which has also been applied in machine translation (Dorr, 1993, 1992).
Computational approaches to the interpretation of metaphor include convention-based and reasoning-based approaches. Convention-based approaches encode specific knowledge about a relatively small core set of conventional metaphors. These represen-
tations are then used during understanding to replace one meaning with an appropriate metaphorical one (Norvig, 1987; Martin, 1990; Hayes and Bayer, 1991; Veale and Keane, 1992; Jones and McCoy, 1992). Reasoning-based approaches eschew representing metaphoric conventions, instead modeling figurative language processing via general reasoning ability, such as analogical reasoning, rather than as a specifically language-related phenomenon. (Russell, 1976; Carbonell, 1982; Gentner, 1983; Fass, 1988, 1991, 1997).
An influential collection of papers on metaphor can be found in Ortony (1993). Lakoff and Johnson (1980) is the classic work on conceptual metaphor and metonymy. Russell (1976) presents one of the earliest computational approaches to metaphor. Additional early work can be found in DeJong and Waltz (1983), Wilks (1978) and Hobbs (1979). More recent computational efforts to analyze metaphor can be found in Fass (1988, 1991, 1997), Martin (1990), Veale and Keane (1992), Iverson and Helmreich (1992), and Chandler (1991). Martin (1996) presents a survey of computational approaches to metaphor and other types of figurative language.
STILL NEEDS SOME UPDATES.
EXERCISES
19.1 Collect three definitions of ordinary non-technical English words from a dictionary of your choice that you feel are flawed in some way. Explain the nature of the flaw and how it might be remedied.
19.2 Give a detailed account of similarities and differences among the following set of lexemes: imitation, synthetic, artificial, fake, and simulated.
19.3 Examine the entries for these lexemes in WordNet (or some dictionary of your choice). How well does it reflect your analysis?
19.4 The WordNet entry for the noun bat lists 6 distinct senses. Cluster these senses using the definitions of homonymy and polysemy given in this chapter. For any senses that are polysemous, give an argument as to how the senses are related.
19.5 Assign the various verb arguments in the following WSJ examples to their appropriate thematic roles using the set of roles shown in Figure 19.6.
a. The intense heat buckled the highway about three feet.
b. He melted her reserve with a husky-voiced paean to her eyes.
c. But Mingo, a major Union Pacific shipping center in the 1890s, has melted away to little more than the grain elevator now.
19.6 Using WordNet, describe appropriate selectional restrictions on the verbs drink, kiss, and write.
19.7 Collect a small corpus of examples of the verbs drink, kiss, and write, and analyze how well your selectional restrictions worked.
19.8 Consider the following examples from (McCawley, 1968):
My neighbor is a father of three.
?My buxom neighbor is a father of three.
What does the ill-formedness of the second example imply about how constituents satisfy, or violate, selectional restrictions?
19.9 Find some articles about business, sports, or politics from your daily newspaper. Identify as many uses of conventional metaphors as you can in these articles. How many of the words used to express these metaphors have entries in either WordNet or your favorite dictionary that directly reflect the metaphor.
19.10 Consider the following example:
The stock exchange wouldn't talk publicly, but a spokesman said a news conference is set for today to introduce a new technology product.
Assuming that stock exchanges are not the kinds of things that can literally talk, give a sensible account for this phrase in terms of a metaphor or metonymy.
19.11 Choose an English verb that occurs in both FrameNet and PropBank. Compare and contrast the FrameNet and PropBank representations of the arguments of the verb.
Amsler, R. A. (1980). The Structure of the Merriam-Webster Pocket Dictionary. Ph.D. thesis, University of Texas, Austin, Texas. Report No.
Amsler, R. A. (1981). A taxonomy of English nouns and verbs. In ACL-81, Stanford, CA, pp. 133–138. ACL.
Baker, C. F., Fillmore, C. J., and Lowe, J. B. (1998). The Berkeley FrameNet project. In COLING/ACL-98, pp. 86–90.
Boguraev, B. and Briscoe, T. (Eds.). (1989). Computational Lexicography for Natural Language Processing. Longman, London.
Carbonell, J. (1982). Metaphor: An inescapable phenomenon in natural language comprehension. In Lehnert, W. G. and Ringle, M. (Eds.), Strategies for Natural Language Processing, pp. 415–434. Lawrence Erlbaum.
Chandler, S. (1991). Metaphor comprehension: A connectionist approach to implications for the mental lexicon. Metaphor and Symbolic Activity, 6(4), 227–258.
Copestake, A. and Briscoe, T. (1995). Semi-productive polysemy and sense extension. Journal of Semantics, 12(1), 15–68.
Cruse, D. A. (2004). Meaning in Language: an Introduction to Semantics and Pragmatics. Oxford University Press, Oxford. Second edition.
DeJong, G. F. and Waltz, D. L. (1983). Understanding novel language. Computers and Mathematics with Applications, 9.
Dorr, B. (1992). The use of lexical semantics in interlingual machine translation. Journal of Machine Translation, 7(3), 135–193.
Dorr, B. (1993). Machine Translation. MIT Press.
Dowty, D. R. (1979). Word Meaning and Montague Grammar. D. Reidel, Dordrecht.
Fass, D. (1988). Collate Semantics: A Semantics for Natural Language. Ph.D. thesis, New Mexico State University, Las Cruces, New Mexico. CRL Report No. MCCS-88-118.
Fass, D. (1991). met*: A method for discriminating metaphor and metonymy by computer. Computational Linguistics, 17(1).
Fass, D. (1997). Processing Metonymy and Metaphor. Ablex Publishing, Greenwich, CT.
Fellbaum, C. (Ed.). (1998). WordNet: An Electronic Lexical Database. MIT Press.
Fillmore, C. J. (1968). The case for case. In Bach, E. W. and Harms, R. T. (Eds.), Universals in Linguistic Theory, pp. 1–88. Holt, Rinehart & Winston.
Fillmore, C. J. (1985). Frames and the semantics of understanding. Quaderni di Semantica, VI(2), 222–254.
Gentner, D. (1983). Structure mapping: A theoretical framework for analogy. Cognitive Science, 7, 155–170.
Hayes, E. and Bayer, S. (1991). Metaphoric generalization through sort coercion. In Proceedings of the 29th ACL, Berkeley, CA, pp. 222–228. ACL.
Gruber, J. S. (1965). Studies in Lexical Relations. Ph.D. thesis, MIT.
Hobbs, J. R. (1979). Metaphor, metaphor schemata, and selective inferencing. Tech. rep. Technical Note 204, SRI.
Iverson, E. and Helmreich, S. (1992). Metalle: An integrated approach to non-literal phrase interpretation. Computational Intelligence, 8(3).
Jackendoff, R. (1983). Semantics and Cognition. MIT Press.
Jackendoff, R. (1990). Semantic Structures. MIT Press.
Johnson-Laird, P. N. (1983). Mental Models. Harvard University Press, Cambridge, MA.
Jones, M. A. and McCoy, K. (1992). Transparently-motivated metaphor generation. In Dale, R., Hovy, E. H., Rösner, D., and Stock, O. (Eds.), Aspects of Automated Natural Language Generation, Lecture Notes in Artificial Intelligence 587, pp. 183–198. Springer Verlag, Berlin.
Katz, J. J. and Fodor, J. A. (1963). The structure of a semantic theory. Language, 39, 170–210.
Kipper, K., Dang, H. T., and Palmer, M. (2000). Class-based construction of a verb lexicon. In Proceedings of the Seventh National Conference on Artificial Intelligence (AAAI-2000), Austin, TX.
Klavans, J. L. (Ed.). (1995). Representation and Acquisition of Lexical Knowledge: Polysemy, Ambiguity and Generativity. AAAI Press, Menlo Park, CA. AAAI Technical Report SS-95-01.
Lakoff, G. (1965). On the Nature of Syntactic Irregularity. Ph.D. thesis, Indiana University. Published as Irregularity in Syntax. Holt, Rinehart, and Winston, New York, 1970.
Lakoff, G. and Johnson, M. (1980). Metaphors We Live By. University of Chicago Press, Chicago, IL.
Lehrer, A. (1974). Semantic Fields and Lexical Structure. North-Holland, Amsterdam.
Lehrer, A. and Kittay, E. (Eds.). (1992). Frames, Fields and Contrasts: New Essays in Semantic and Lexical Organization. Lawrence Erlbaum.
Levin, B. (1993). English Verb Classes And Alternations: A Preliminary Investigation. University of Chicago Press, Chicago.
Levin, B. and Rappaport Hovav, M. (2005). Argument Realization. Cambridge University Press.
Lowe, J. B., Baker, C. F., and Fillmore, C. J. (1997). A frame-semantic approach to semantic annotation. In Proceedings of ACL SIGLEX Workshop on Tagging Text with Lexical Semantics, Washington, D.C., pp. 18–24. ACL.
Lyons, J. (1977). Semantics. Cambridge University Press.
Martin, J. H. (1986). The acquisition of polysemy. In ICML 1986, Irvine, CA, pp. 198–204.
Martin, J. H. (1990). A Computational Model of Metaphor Interpretation. Perspectives in Artificial Intelligence. Academic Press.
Martin, J. H. (1996). Computational approaches to figurative language. Metaphor and Symbolic Activity, 11(1), 85–100.
McCawley, J. D. (1968). The role of semantics in a grammar. In Bach, E. W. and Harms, R. T. (Eds.), Universals in Linguistic Theory, pp. 124–169. Holt, Rinehart & Winston.
Morris, W. (Ed.). (1985). American Heritage Dictionary (2nd College Edition edition). Houghton Mifflin.
Nida, E. A. (1975). Componential Analysis of Meaning: An Introduction to Semantic Structures. Mouton, The Hague.
Norvig, P. (1987). A Unified Theory of Inference for Text Understanding. Ph.D. thesis, University of California, Berkeley, CA. Available as University of California at Berkeley Computer Science Division Tech. rep. #87/339.
Ortony, A. (Ed.). (1993). Metaphor (2nd edition). Cambridge University Press, Cambridge.
Pustejovsky, J. (1995). The Generative Lexicon. MIT Press.
Pustejovsky, J. and Bergler, S. (Eds.). (1992). Lexical Semantics and Knowledge Representation. Lecture Notes in Artificial Intelligence. Springer Verlag, Berlin.
Pustejovsky, J. and Boguraev, B. (Eds.). (1996). Lexical Semantics: The Problem of Polysemy. Oxford University Press, Oxford.
Ruppenhofer, J., Ellsworth, M., Petruck, M. R. L., Johnson, C. R., and Scheffczyk, J. (2006). FrameNet ii: Extended theory and practice. Version 1.3, http://www.icsi.berkeley.edu/framenet/.
Russell, S. W. (1976). Computer understanding of metaphorically used verbs. American Journal of Computational Linguistics, 2. Microfiche 44.
Saint-Dizier, P. and Viegas, E. (Eds.). (1995). Computational Lexical Semantics. Cambridge University Press.
Schank, R. C. and Albelson, R. P. (1977). Scripts, Plans, Goals and Understanding. Lawrence Erlbaum.
Simmons, R. F. (1973). Semantic networks: Their computation and use for understanding English sentences. In Schank, R. C. and Colby, K. M. (Eds.), Computer Models of Thought and Language, pp. 61–113. W.H. Freeman and Co., San Francisco.
Simmons, R. F. (1978). Rule-based computations on English. In Waterman, D. A. and Hayes-Roth, F. (Eds.), Pattern-Directed Inference Systems. Academic Press.
Simmons, R. F. (1983). Computations from the English. Prentice Hall.
Veale, T. and Keane, M. T. (1992). Conceptual scaffolding: A spatially founded meaning representation for metaphor comprehension. Computational Intelligence, 8(3), 494–519.
Wierzbicka, A. (1992). Semantics, Culture, and Cognition: University Human Concepts in Culture-Specific Configurations. Oxford University Press.
Wierzbicka, A. (1996). Semantics: Primes and Universals. Oxford University Press.
Wilks, Y. (1975). An intelligent analyzer and understander of English. Communications of the ACM, 18(5), 264–274.
Wilks, Y. (1978). Making preferences more active. Artificial Intelligence, 11(3), 197–223.
20 COMPUTATIONAL LEXICAL SEMANTICS
To get a single right meaning is better than a ship-load of pearls. To resolve a single doubt is like the bottom falling off the bucket.
Yuen Mei (1785) (translation by Arthur Waley)
The asphalt that Los Angeles is famous for occurs mainly on its freeways. But in the middle of the city is another patch of asphalt, the La Brea tar pits, and this asphalt preserves millions of fossil bones from the last of the Ice Ages of the Pleistocene Epoch. One of these fossils is the Smilodon, or sabre-toothed tiger, instantly recognizable by its long canines. Five million years ago or so, a completely different sabre-tooth tiger called Thylacosmilus lived in Argentina and other parts of South America. Thylacosmilus was a marsupial where Smilodon was a placental mammal, but had the same long upper canines and, like Smilodon, had a protective bone flange on the lower jaw. The similarity of these two mammals is one of many examples of parallel or convergent evolution, in which particular contexts or environments lead to the evolution of very similar structures in different species (Gould, 1980).
The role of context is also important in the similarity of a less biological kind of organism: the word. Suppose we wanted to decide if two words have similar meanings. Not surprisingly, words with similar meanings often occur in similar contexts, whether in terms of corpora (having similar neighboring words or syntactic structures in sentences) or in terms of dictionaries and thesauruses (having similar definitions, or being nearby in the thesaurus hierarchy). Thus similarity of context turns out to be an important way to detect semantic similarity. Semantic similarity turns out to play an important role in a diverse set of applications including information retrieval, question answering, summarization and generation, text classification, automatic essay grading and the detection of plagiarism.
In this chapter we introduce a series of topics related to computing with word meanings, or computational lexical semantics. Roughly in parallel with the sequence of topics in Ch. 19, we'll introduce computational tasks associated with word senses, relations among words, and the thematic structure of predicate-bearing words. We'll see the role of important role of context and similarity of sense in each of these.
We begin with word sense disambiguation, the task of examining word tokens in context and determining which sense of each word is being used. WSD is a task with
a long history in computational linguistics, and as we will see, is a non-trivial undertaking given the somewhat elusive nature of many word senses. Nevertheless, there are robust algorithms that can achieve high levels of accuracy given certain reasonable assumptions. Many of these algorithms rely on contextual similarity to help choose the proper sense.
This will lead us natural to a consideration of the computation of word similarity and other relations between words, including the hypernym, hyponym, and meronym WordNet relations introduced in Ch. 19. We'll introduce methods based purely on corpus similarity, and others based on structured resources such as WordNet.
Finally, we describe algorithms for semantic role labeling, also known as case role or thematic role assignment. These algorithms generally use features extracted from syntactic parses to assign semantic roles such as AGENT, THEME and INSTRUMENT to the phrases in a sentence with respect to particular predicates.
20.1 WORD SENSE DISAMBIGUATION: OVERVIEW
Our discussion of compositional semantic analyzers in Ch. 18 pretty much ignored the issue of lexical ambiguity. It should be clear by now that this is an unreasonable approach. Without some means of selecting correct senses for the words in an input, the enormous amount of homonymy and polysemy in the lexicon would quickly overwhelm any approach in an avalanche of competing interpretations.
The task of selecting the correct sense for a word is called word sense disambiguation, or WSD. Disambiguating word senses has the potential to improve many natural language processing tasks. As we'll see in Ch. 25, machine translation is one area where word sense ambiguities can cause severe problems; others include question-answering, information retrieval, and text classification. The way that WSD is exploited in these and other applications varies widely based on the particular needs of the application. The discussion presented here ignores these application-specific differences and focuses on the implementation and evaluation of WSD systems as a stand-alone task.
In their most basic form, WSD algorithms take as input a word in context along with a fixed inventory of potential word senses, and return the correct word sense for that use. Both the nature of the input and the inventory of senses depends on the task. For machine translation from English to Spanish, the sense tag inventory for an English word might be the set of different Spanish translations. If speech synthesis is our task, the inventory might be restricted to homographs with differing pronunciations such as bass and bow. If our task is automatic indexing of medical articles, the sense tag inventory might be the set of MeSH (Medical Subject Headings) thesaurus entries. When we are evaluating WSD in isolation, we can use the set of senses from a dictionary/thesaurus resource like WordNet or LDOCE. Fig. 20.1 shows an example for the word bass, which can refer to a musical instrument or a kind of fish. $ ^{1} $
| WordNet Sense | Spanish Translation | Roget Category | Target Word in Context |
| --- | --- | --- | --- |
| $ bass^{4} $ | lubina | FISH/INSECT | ...fish as Pacific salmon and striped bass and... |
| $ bass^{4} $ | lubina | FISH/INSECT | ...produce filets of smoked bass or sturgeon... |
| $ bass^{7} $ | bajo | MUSIC | ...exciting jazz bass player since Ray Brown... |
| $ bass^{7} $ | bajo | MUSIC | ...play bass because he doesn't have to solo... |
It is useful to distinguish two variants of the generic WSD task. In the lexical sample task, a small pre-selected set of target words is chosen, along with an inventory of senses for each word from some lexicon. Since the set of words and the set of senses is small, supervised machine learning approaches are often used to handle lexical sample tasks. For each word, a number of corpus instances (context sentences) can be selected and hand-labeled with the correct sense of the target word in each. Classifier systems can then be trained using these labeled examples. Unlabeled target words in context can then be labeled using such a trained classifier. Early work in word sense disambiguation focused solely on lexical sample tasks of this sort, building word-specific algorithms for disambiguating single words like line, interest, or plant.
In contrast, in the all-words task systems are given entire texts and a lexicon with an inventory of senses for each entry, and are required to disambiguate every content word in the text. The all-words task is very similar to part-of-speech tagging, except with a much larger set of tags, since each lemma has its own set. A consequence of this larger set of tags is a serious data sparseness problem; there is unlikely to be adequate training data for every word in the test set. Moreover, given the number of polysemous words in reasonably-sized lexicons, approaches based on training one classifier per term are unlikely to be practical.
In the following sections we explore the application of various machine learning paradigms to word sense disambiguation. We begin with supervised learning, followed by a section on how systems are standardly evaluated. We then turn to a variety of methods for dealing with the lack of sufficient day for fully-supervised training, including dictionary-based approaches and bootstrapping techniques.
Finally, after we have introduced the necessary notions of distributional word similarity in Sec. 20.7, we return in Sec. 20.10 to the problem of unsupervised approaches to sense disambiguation.
20.2 SUPERVISED WORD SENSE DISAMBIGUATION
If we have data which has been hand-labeled with correct word senses, we can use a supervised learning approach to the problem of sense disambiguation. Extracting features from the text that are helpful in predicting particular senses, and then training a classifier to assign the correct sense given these features. The output of training is thus a classifier system capable of assigning sense labels to unlabeled words in context.
For lexical sample tasks, there are various labeled corpora for individual words, consisting of context sentences labeled with the correct sense for the target word. These
include the line-hard-serve corpus containing 4,000 sense-tagged examples of line as a noun, hard as an adjective and serve as a verb (Leacock et al., 1993), and the interest corpus with 2,369 sense-tagged examples of interest as a noun (Bruce and Wiebe, 1994). The SENSEVAL project has also produced a number of such sense-labeled lexical sample corpora (SENSEVAL-1 with 34 words from the HECTOR lexicon and corpus (Kilgarriff and Rosenzweig, 2000; Atkins, 1993), SENSEVAL-2 and -3 with 73 and 57 target words, respectively (Palmer et al., 2001; Kilgarriff, 2001)).
For training all-word disambiguation tasks we use a semantic concordance, a corpus in which each open-class word in each sentence is labeled with its word sense from a specific dictionary or thesaurus. One commonly used corpus is SemCor, a subset of the Brown Corpus consisting of over 234,000 words which were manually tagged with WordNet senses (Miller et al., 1993; Landes et al., 1998). In addition, sense-tagged corpora have been built for the SENSEVAL all-word tasks. The SENSEVAL-3 English all-words test data consisted of 2081 tagged content word tokens, from 5,000 total running words of English from the WSJ and Brown corpora (Palmer et al., 2001).