12.4.2 Using a Treebank as a Grammar
The sentences in a treebank implicitly constitute a grammar of the language. For example, we can take the three parsed sentences in Fig. 12.7 and Fig. 12.9 and extract each of the CFG rules in them. For simplicity, let's strip off the rule suffixes (-SBJ and so on). The resulting grammar is shown in Fig. 12.10.
The grammar used to parse the Penn Treebank is relatively flat, resulting in very many and very long rules. For example among the approximately 4,500 different rules for expanding VP are separate rules for PP sequences of any length, and every possible arrangement of verb arguments:
VP → VBD PP
VP → VBD PP PP
VP → VBD PP PP PP
VP → VBD PP PP PP PP
VP $ \rightarrow $ VB PP ADVP
VP $ \rightarrow $ ADVP VB PP
as well as even longer rules, such a:
VP $ \rightarrow $ VBP PP PP PP PP PP PP ADVP PP
which comes from the VP marked in italics:
(12.7) This mostly happens because we go from football in the fall to lifting in the winter to football again in the spring.
Some of the many thousands of NP rules include:
NP → DT JJ NN
NP → DT JJ NNS
NP → DT JJ NN NN
NP → DT JJ JJ NN
NP → DT JJ CD NNS
NP → RB DT JJ NN NN
NP → RB DT JJ JJ NNS
NP → DT JJ JJ NNP NNS
NP → DT NNP NNP NNP NNP JJ NN
NP → DT JJ NNP CC JJ JJ NN NNS
NP → RB DT JJS NN NN SBAR
NP → DT VBG JJ NNP NNP CC NNP
NP → DT JJ NNS, NNS CC NN NNS NN
NP → DT JJ VBG NN NNP NNP FW NNP
NP → NP JJ, JJ ' SBAR ' NNS
The last two of those rules, for example, come from the following two NPs:
(12.8) $ [_{DT} The] $ [JJ state-owned] [JJ industrial] [VBG holding] [NN company] [NNP Instituto] [NNP Nacional] [FW de] [NNP Industria]
(12.9) $ [_{NP} Shearson's] $ [JJ easy-to-film], [JJ black-and-white] "[SBAR Where We Stand]" [NNS commercials]
Viewed as a large grammar in this way, the Penn Treebank III Wall Street Journal corpus, which contains about 1 million words, also has about 1 million non-lexical rule tokens, consisting of about 17,500 distinct rule types.
Various facts about the treebank grammars, such as their large numbers of flat rules, pose problems for probabilistic parsing algorithms. For this reason, it is common to make various modifications to a grammar extracted from a treebank. We will discuss these further in Ch. 14.