← 学习库 Introductory Statistics (OpenStax) · 中英对照 目录

2 Descriptive Statistics 描述统计学

本页译自 OpenStax《Introductory Statistics》第 2 章 Descriptive Statistics。公式经本地 MathJax 渲染,自定义宏已注入。

Introduction 引言

Once you have collected data, what will you do with it? Data can be described and presented in many different formats. For example, suppose you are interested in buying a house in a particular area. You may have no clue about the house prices, so you might ask your real estate agent to give you a sample data set of prices. Looking at all the prices in the sample often is overwhelming. A better way might be to look at the median price and the variation of prices. The median and variation are just two ways that you will learn to describe data. Your agent might also provide you with a graph of the data.

收集到数据后,你打算如何处理?数据可以用许多不同的格式来描述和呈现。例如,假设你有意在某个特定区域买房。你可能对房价一无所知,因而会请房产中介给你一份价格的样本数据集。逐一查看样本中的所有价格往往令人不知所措。更好的办法也许是关注价格的中位数和价格的离散程度。中位数和离散程度只是你将要学习的描述数据的两种方式。你的中介还可能会给你提供一份数据的图形。

In this chapter, you will study numerical and graphical ways to describe and display your data. This area of statistics is called "Descriptive Statistics." You will learn how to calculate, and even more importantly, how to interpret these measurements and graphs.

在本章中,你将学习用数值和图形的方法来描述和展示你的数据。统计学的这一领域称为"描述统计学"。你将学习如何计算——更重要的是——如何解释这些度量和图形。

A statistical graph is a tool that helps you learn about the shape or distribution of a sample or a population. A graph can be a more effective way of presenting data than a mass of numbers because we can see where data clusters and where there are only a few data values. Newspapers and the Internet use graphs to show trends and to enable readers to compare facts and figures quickly. Statisticians often graph data first to get a picture of the data. Then, more formal tools may be applied.

统计图形是一种帮助你了解样本或总体形态或分布的工具。与一大堆数字相比,图形是展示数据更有效的方法,因为我们能从中看出数据聚集在何处、何处只有少数数据值。报纸和互联网使用图形来显示趋势,并让读者快速比较事实与数字。统计学家通常先对数据作图以获得数据的概貌,然后再应用更正规的方法。

Some of the types of graphs that are used to summarize and organize data are the dot plot, the bar graph, the histogram, the stem-and-leaf plot, the frequency polygon (a type of broken line graph), the pie chart, and the box plot. In this chapter, we will briefly look at stem-and-leaf plots, line graphs, and bar graphs, as well as frequency polygons, and time series graphs. Our emphasis will be on histograms and box plots.

用于概括和组织数据的图形类型包括点图、条形图、直方图、茎叶图、频数多边形(一种折线图)、饼图和箱线图。在本章中,我们将简要介绍茎叶图、线图和条形图,以及频数多边形和时间序列图。我们的重点是直方图和箱线图。

2.1 Stem-and-Leaf Graphs (Stemplots), Line Graphs, and Bar Graphs 2.1 茎叶图(Stemplots)、线图与条形图

One simple graph, the stem-and-leaf graph or stemplot, comes from the field of exploratory data analysis. It is a good choice when the data sets are small. To create the plot, divide each observation of data into a stem and a leaf. The leaf consists of a final significant digit. For example, 23 has stem two and leaf three. The number 432 has stem 43 and leaf two. Likewise, the number 5,432 has stem 543 and leaf two. The decimal 9.3 has stem nine and leaf three. Write the stems in a vertical line from smallest to largest. Draw a vertical line to the right of the stems. Then write the leaves in increasing order next to their corresponding stem.

一种简单的图形——茎叶图(stem-and-leaf graph)或茎叶图(stemplot)——源自探索性数据分析领域。当数据集较小时,它是很好的选择。要绘制该图,需将每个数据观测值分成"茎"和"叶"两部分。叶由最后一位有效数字组成。例如,23 的茎是 2、叶是 3;432 的茎是 43、叶是 2;同理,5432 的茎是 543、叶是 2;小数 9.3 的茎是 9、叶是 3。将茎按从小到大的顺序纵向排列,在茎的右侧画一条竖线,然后按递增顺序把叶写在对应的茎旁边。

For Susan Dean's spring pre-calculus class, scores for the first exam were as follows (smallest to largest):

在 Susan Dean 春季的初等微积分课上,第一次考试的成绩如下(从小到大):

33; 42; 49; 49; 53; 55; 55; 61; 63; 67; 68; 68; 69; 69; 72; 73; 74; 78; 80; 83; 88; 88; 88; 90; 92; 94; 94; 94; 94; 96; 100

33;42;49;49;53;55;55;61;63;67;68;68;69;69;72;73;74;78;80;83;88;88;88;90;92;94;94;94;94;96;100

| Stem | Leaf |

| 茎 | 叶 |

|------|---------------|

|------|---------------|

| 3 | 3 |

| 3 | 3 |

| 4 | 2 9 9 |

| 4 | 2 9 9 |

| 5 | 3 5 5 |

| 5 | 3 5 5 |

| 6 | 1 3 7 8 8 9 9 |

| 6 | 1 3 7 8 8 9 9 |

| 7 | 2 3 4 8 |

| 7 | 2 3 4 8 |

| 8 | 0 3 8 8 8 |

| 8 | 0 3 8 8 8 |

| 9 | 0 2 4 4 4 4 6 |

| 9 | 0 2 4 4 4 4 6 |

| 10 | 0 |

| 10 | 0 |

Table 2.1 Stem-and-Leaf Graph

表 2.1 茎叶图

The stemplot shows that most scores fell in the 60s, 70s, 80s, and 90s. Eight out of the 31 scores or approximately 26% $\left( \frac{8}{31} \right)$ were in the 90s or 100, a fairly high number of As.

茎叶图显示,大多数分数落在 60、70、80 和 90 多分段。31 个分数中有 8 个(约为 26% $\left( \frac{8}{31} \right)$)在 90 分或以上,获得 A 的人数相当多。

For the Park City basketball team, scores for the last 30 games were as follows (smallest to largest):

在 Park City 篮球队,最近 30 场比赛的得分如下(从小到大):

32; 32; 33; 34; 38; 40; 42; 42; 43; 44; 46; 47; 47; 48; 48; 48; 49; 50; 50; 51; 52; 52; 52; 53; 54; 56; 57; 57; 60; 61

32;32;33;34;38;40;42;42;43;44;46;47;47;48;48;48;49;50;50;51;52;52;52;53;54;56;57;57;60;61

Construct a stem plot for the data.

为这些数据绘制一张茎叶图。

The stemplot is a quick way to graph data and gives an exact picture of the data. You want to look for an overall pattern and any outliers. An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500) while others may indicate that something unusual is happening. It takes some background information to explain outliers, so we will cover them in more detail later.

茎叶图是绘制数据的快捷方式,并能精确呈现数据。你需要寻找整体的模式和任何离群值。离群值是不符合其余数据的观测值,有时也称为极值。当你作图时,离群值会显得不符合图形的模式。有些离群值源于错误(例如把 500 写成了 50),而另一些可能表明发生了异常情况。解释离群值需要一些背景信息,因此我们稍后会更详细地讨论它们。

The data are the distances (in kilometers) from a home to local supermarkets. Create a stemplot using the data:

以下是某住所到本地超市的距离(单位:千米)。用这些数据绘制一张茎叶图:

1.1; 1.5; 2.3; 2.5; 2.7; 3.2; 3.3; 3.3; 3.5; 3.8; 4.0; 4.2; 4.5; 4.5; 4.7; 4.8; 5.5; 5.6; 6.5; 6.7; 12.3

1.1;1.5;2.3;2.5;2.7;3.2;3.3;3.3;3.5;3.8;4.0;4.2;4.5;4.5;4.7;4.8;5.5;5.6;6.5;6.7;12.3

Problem 问题

Do the data seem to have any concentration of values?

这些数据似乎在某些数值上有所集中吗?

The leaves are to the right of the decimal.

叶位于小数点右侧。

Solution 解答

The value 12.3 may be an outlier. Values appear to concentrate at three and four kilometers.

数值 12.3 可能是一个离群值。数值似乎集中在 3 千米和 4 千米附近。

| Stem | Leaf |

| 茎 | 叶 |

|------|-------------|

|------|-------------|

| 1 | 1 5 |

| 1 | 1 5 |

| 2 | 3 5 7 |

| 2 | 3 5 7 |

| 3 | 2 3 3 5 8 |

| 3 | 2 3 3 5 8 |

| 4 | 0 2 5 5 7 8 |

| 4 | 0 2 5 5 7 8 |

| 5 | 5 6 |

| 5 | 5 6 |

| 6 | 5 7 |

| 6 | 5 7 |

| 7 | |

| 7 | |

| 8 | |

| 8 | |

| 9 | |

| 9 | |

| 10 | |

| 10 | |

| 11 | |

| 11 | |

| 12 | 3 |

| 12 | 3 |

Table 2.2

表 2.2

The following data show the distances (in miles) from the homes of off-campus statistics students to the college. Create a stem plot using the data and identify any outliers:

以下数据是该学院校外统计学学生住所到学院的距离(单位:英里)。用这些数据绘制一张茎叶图,并指出任何离群值:

0.5; 0.7; 1.1; 1.2; 1.2; 1.3; 1.3; 1.5; 1.5; 1.7; 1.7; 1.8; 1.9; 2.0; 2.2; 2.5; 2.6; 2.8; 2.8; 2.8; 3.5; 3.8; 4.4; 4.8; 4.9; 5.2; 5.5; 5.7; 5.8; 8.0

0.5;0.7;1.1;1.2;1.2;1.3;1.3;1.5;1.5;1.7;1.7;1.8;1.9;2.0;2.2;2.5;2.6;2.8;2.8;2.8;3.5;3.8;4.4;4.8;4.9;5.2;5.5;5.7;5.8;8.0

Problem 问题

A side-by-side stem-and-leaf plot allows a comparison of the two data sets in two columns. In a side-by-side stem-and-leaf plot, two sets of leaves share the same stem. The leaves are to the left and the right of the stems. Table 2.4 and Table 2.5 show the ages of presidents at their inauguration and at their death. Construct a side-by-side stem-and-leaf plot using this data.

并列茎叶图(side-by-side stem-and-leaf plot)可以对两列中的两个数据集进行比较。在并列茎叶图中,两组叶共享同一个茎,叶分布在茎的左侧和右侧。表 2.4 和表 2.5 显示了各位总统在就职时和去世时的年龄。用这些数据绘制一张并列茎叶图。

Solution 解答

| Ages at Inauguration | | Ages at Death |

| 就职时年龄 | | 去世时年龄 |

|---------------------------------------------------|-----|-------------------------|

|---------------------------------------------------|-----|-------------------------|

| 9 9 8 7 7 7 6 3 2 | 4 | 6 9 |

| 9 9 8 7 7 7 6 3 2 | 4 | 6 9 |

| 8 7 7 7 7 6 6 6 5 5 5 5 4 4 4 4 4 2 2 1 1 1 1 1 0 | 5 | 3 6 6 7 7 8 |

| 8 7 7 7 7 6 6 6 5 5 5 5 4 4 4 4 4 2 2 1 1 1 1 1 0 | 5 | 3 6 6 7 7 8 |

| 9 8 5 4 4 2 1 1 1 0 | 6 | 0 0 3 3 4 4 5 6 7 7 7 8 |

| 9 8 5 4 4 2 1 1 1 0 | 6 | 0 0 3 3 4 4 5 6 7 7 7 8 |

| | 7 | 0 1 1 1 2 3 4 7 8 8 9 |

| | 7 | 0 1 1 1 2 3 4 7 8 8 9 |

| | 8 | 0 1 3 5 8 |

| | 8 | 0 1 3 5 8 |

| | 9 | 0 0 3 3 |

| | 9 | 0 0 3 3 |

Table 2.3

表 2.3

| President | Age | President | Age | President | Age |

| 总统 | 年龄 | 总统 | 年龄 | 总统 | 年龄 |

|----------------|-----|--------------|-----|--------------|-----|

|----------------|-----|--------------|-----|--------------|-----|

| Washington | 57 | Lincoln | 52 | Hoover | 54 |

| Washington | 57 | Lincoln | 52 | Hoover | 54 |

| J. Adams | 61 | A. Johnson | 56 | F. Roosevelt | 51 |

| J. Adams | 61 | A. Johnson | 56 | F. Roosevelt | 51 |

| Jefferson | 57 | Grant | 46 | Truman | 60 |

| Jefferson | 57 | Grant | 46 | Truman | 60 |

| Madison | 57 | Hayes | 54 | Eisenhower | 62 |

| Madison | 57 | Hayes | 54 | Eisenhower | 62 |

| Monroe | 58 | Garfield | 49 | Kennedy | 43 |

| Monroe | 58 | Garfield | 49 | Kennedy | 43 |

| J. Q. Adams | 57 | Arthur | 51 | L. Johnson | 55 |

| J. Q. Adams | 57 | Arthur | 51 | L. Johnson | 55 |

| Jackson | 61 | Cleveland | 47 | Nixon | 56 |

| Jackson | 61 | Cleveland | 47 | Nixon | 56 |

| Van Buren | 54 | B. Harrison | 55 | Ford | 61 |

| Van Buren | 54 | B. Harrison | 55 | Ford | 61 |

| W. H. Harrison | 68 | Cleveland | 55 | Carter | 52 |

| W. H. Harrison | 68 | Cleveland | 55 | Carter | 52 |

| Tyler | 51 | McKinley | 54 | Reagan | 69 |

| Tyler | 51 | McKinley | 54 | Reagan | 69 |

| Polk | 49 | T. Roosevelt | 42 | G.H.W. Bush | 64 |

| Polk | 49 | T. Roosevelt | 42 | G.H.W. Bush | 64 |

| Taylor | 64 | Taft | 51 | Clinton | 47 |

| Taylor | 64 | Taft | 51 | Clinton | 47 |

| Fillmore | 50 | Wilson | 56 | G. W. Bush | 54 |

| Fillmore | 50 | Wilson | 56 | G. W. Bush | 54 |

| Pierce | 48 | Harding | 55 | Obama | 47 |

| Pierce | 48 | Harding | 55 | Obama | 47 |

| Buchanan | 65 | Coolidge | 51 | | |

| Buchanan | 65 | Coolidge | 51 | | |

Table 2.4 Presidential Ages at Inauguration

表 2.4 总统就职时年龄

| President | Age | President | Age | President | Age |

| 总统 | 年龄 | 总统 | 年龄 | 总统 | 年龄 |

|----------------|-----|--------------|-----|--------------|-----|

|----------------|-----|--------------|-----|--------------|-----|

| Washington | 67 | Lincoln | 56 | Hoover | 90 |

| Washington | 67 | Lincoln | 56 | Hoover | 90 |

| J. Adams | 90 | A. Johnson | 66 | F. Roosevelt | 63 |

| J. Adams | 90 | A. Johnson | 66 | F. Roosevelt | 63 |

| Jefferson | 83 | Grant | 63 | Truman | 88 |

| Jefferson | 83 | Grant | 63 | Truman | 88 |

| Madison | 85 | Hayes | 70 | Eisenhower | 78 |

| Madison | 85 | Hayes | 70 | Eisenhower | 78 |

| Monroe | 73 | Garfield | 49 | Kennedy | 46 |

| Monroe | 73 | Garfield | 49 | Kennedy | 46 |

| J. Q. Adams | 80 | Arthur | 56 | L. Johnson | 64 |

| J. Q. Adams | 80 | Arthur | 56 | L. Johnson | 64 |

| Jackson | 78 | Cleveland | 71 | Nixon | 81 |

| Jackson | 78 | Cleveland | 71 | Nixon | 81 |

| Van Buren | 79 | B. Harrison | 67 | Ford | 93 |

| Van Buren | 79 | B. Harrison | 67 | Ford | 93 |

| W. H. Harrison | 68 | Cleveland | 71 | Reagan | 93 |

| W. H. Harrison | 68 | Cleveland | 71 | Reagan | 93 |

| Tyler | 71 | McKinley | 58 | | |

| Tyler | 71 | McKinley | 58 | | |

| Polk | 53 | T. Roosevelt | 60 | | |

| Polk | 53 | T. Roosevelt | 60 | | |

| Taylor | 65 | Taft | 72 | | |

| Taylor | 65 | Taft | 72 | | |

| Fillmore | 74 | Wilson | 67 | | |

| Fillmore | 74 | Wilson | 67 | | |

| Pierce | 64 | Harding | 57 | | |

| Pierce | 64 | Harding | 57 | | |

| Buchanan | 77 | Coolidge | 60 | | |

| Buchanan | 77 | Coolidge | 60 | | |

Table 2.5 Presidential Age at Death

表 2.5 总统去世时年龄

The table shows the number of wins and losses the Atlanta Hawks have had in 42 seasons. Create a side-by-side stem-and-leaf plot of these wins and losses.

该表显示了 Atlanta Hawks 队在 42 个赛季中的胜场和负场数。为这些胜场和负场绘制一张并列茎叶图。

| Losses | Wins | Year | Losses | Wins | Year |

| 负场 | 胜场 | 赛季 | 负场 | 胜场 | 赛季 |

|--------|------|-----------|--------|------|-----------|

|--------|------|-----------|--------|------|-----------|

| 34 | 48 | 1968–1969 | 41 | 41 | 1989–1990 |

| 34 | 48 | 1968–1969 | 41 | 41 | 1989–1990 |

| 34 | 48 | 1969–1970 | 39 | 43 | 1990–1991 |

| 34 | 48 | 1969–1970 | 39 | 43 | 1990–1991 |

| 46 | 36 | 1970–1971 | 44 | 38 | 1991–1992 |

| 46 | 36 | 1970–1971 | 44 | 38 | 1991–1992 |

| 46 | 36 | 1971–1972 | 39 | 43 | 1992–1993 |

| 46 | 36 | 1971–1972 | 39 | 43 | 1992–1993 |

| 36 | 46 | 1972–1973 | 25 | 57 | 1993–1994 |

| 36 | 46 | 1972–1973 | 25 | 57 | 1993–1994 |

| 47 | 35 | 1973–1974 | 40 | 42 | 1994–1995 |

| 47 | 35 | 1973–1974 | 40 | 42 | 1994–1995 |

| 51 | 31 | 1974–1975 | 36 | 46 | 1995–1996 |

| 51 | 31 | 1974–1975 | 36 | 46 | 1995–1996 |

| 53 | 29 | 1975–1976 | 26 | 56 | 1996–1997 |

| 53 | 29 | 1975–1976 | 26 | 56 | 1996–1997 |

| 51 | 31 | 1976–1977 | 32 | 50 | 1997–1998 |

| 51 | 31 | 1976–1977 | 32 | 50 | 1997–1998 |

| 41 | 41 | 1977–1978 | 19 | 31 | 1998–1999 |

| 41 | 41 | 1977–1978 | 19 | 31 | 1998–1999 |

| 36 | 46 | 1978–1979 | 54 | 28 | 1999–2000 |

| 36 | 46 | 1978–1979 | 54 | 28 | 1999–2000 |

| 32 | 50 | 1979–1980 | 57 | 25 | 2000–2001 |

| 32 | 50 | 1979–1980 | 57 | 25 | 2000–2001 |

| 51 | 31 | 1980–1981 | 49 | 33 | 2001–2002 |

| 51 | 31 | 1980–1981 | 49 | 33 | 2001–2002 |

| 40 | 42 | 1981–1982 | 47 | 35 | 2002–2003 |

| 40 | 42 | 1981–1982 | 47 | 35 | 2002–2003 |

| 39 | 43 | 1982–1983 | 54 | 28 | 2003–2004 |

| 39 | 43 | 1982–1983 | 54 | 28 | 2003–2004 |

| 42 | 40 | 1983–1984 | 69 | 13 | 2004–2005 |

| 42 | 40 | 1983–1984 | 69 | 13 | 2004–2005 |

| 48 | 34 | 1984–1985 | 56 | 26 | 2005–2006 |

| 48 | 34 | 1984–1985 | 56 | 26 | 2005–2006 |

| 32 | 50 | 1985–1986 | 52 | 30 | 2006–2007 |

| 32 | 50 | 1985–1986 | 52 | 30 | 2006–2007 |

| 25 | 57 | 1986–1987 | 45 | 37 | 2007–2008 |

| 25 | 57 | 1986–1987 | 45 | 37 | 2007–2008 |

| 32 | 50 | 1987–1988 | 35 | 47 | 2008–2009 |

| 32 | 50 | 1987–1988 | 35 | 47 | 2008–2009 |

| 30 | 52 | 1988–1989 | 29 | 53 | 2009–2010 |

| 30 | 52 | 1988–1989 | 29 | 53 | 2009–2010 |

Table 2.6

表 2.6

Another type of graph that is useful for specific data values is a line graph. In the particular line graph shown in Example 2.4, the ***x*-axis (horizontal axis) consists of data values and the *y*-axis (vertical axis) consists of frequency points**. The frequency points are connected using line segments.

另一种对特定数据值有用的图形是线图。在例 2.4 所示的那张线图中,*x* 轴(横轴)由数据值构成,*y* 轴(纵轴)由频数点构成。频数点用线段连接起来。

In a survey, 40 mothers were asked how many times per week a teenager must be reminded to do his or her chores. The results are shown in Table 2.7 and in Figure 2.2.

在一项调查中,40 位母亲被问到,她们每周需要提醒青少年做琐事的次数是多少。结果如表 2.7 和图 2.2 所示。

| Number of times teenager is reminded | Frequency |

| 提醒青少年的次数 | 频数 |

|--------------------------------------|-----------|

|--------------------------------------|-----------|

| 0 | 2 |

| 0 | 2 |

| 1 | 5 |

| 1 | 5 |

| 2 | 8 |

| 2 | 8 |

| 3 | 14 |

| 3 | 14 |

| 4 | 7 |

| 4 | 7 |

| 5 | 4 |

| 5 | 4 |

Table 2.7

表 2.7

In a survey, 40 people were asked how many times per year they had their car in the shop for repairs. The results are shown in Table 2.8. Construct a line graph.

在一项调查中,40 人被问到他们每年把车送修多少次。结果如表 2.8 所示。请绘制一张线图。

| Number of times in shop | Frequency |

| 送修次数 | 频数 |

|-------------------------|-----------|

|-------------------------|-----------|

| 0 | 7 |

| 0 | 7 |

| 1 | 10 |

| 1 | 10 |

| 2 | 14 |

| 2 | 14 |

| 3 | 9 |

| 3 | 9 |

Table 2.8

表 2.8

Bar graphs consist of bars that are separated from each other. The bars can be rectangles or they can be rectangular boxes (used in three-dimensional plots), and they can be vertical or horizontal. The bar graph shown in Example 2.5 has age groups represented on the ***x*-axis and proportions on the *y*-axis**.

条形图由彼此分离的条形组成。这些条形可以是矩形,也可以是长方体(用于三维图形),可以是竖直的也可以是水平的。例 2.5 所示的条形图以年龄组表示在 *x* 轴上,以比例表示在 *y* 轴上。

Problem 问题

By the end of 2011, Facebook had over 146 million users in the United States. Table 2.9 shows three age groups, the number of users in each age group, and the proportion (%) of users in each age group. Construct a bar graph using this data.

到 2011 年底,Facebook 在美国拥有超过 1.46 亿用户。表 2.9 显示了三个年龄组、每个年龄组的用户数以及每个年龄组用户所占的比例(%)。请用这些数据绘制一张条形图。

| Age groups | Number of Facebook users | Proportion (%) of Facebook users |

| 年龄组 | Facebook 用户数 | Facebook 用户比例(%) |

|------------|--------------------------|----------------------------------|

|------------|--------------------------|----------------------------------|

| 13–25 | 65,082,280 | 45% |

| 13–25 | 65,082,280 | 45% |

| 26–44 | 53,300,200 | 36% |

| 26–44 | 53,300,200 | 36% |

| 45–64 | 27,885,100 | 19% |

| 45–64 | 27,885,100 | 19% |

Table 2.9

表 2.9

Solution 解答

The population in Park City is made up of children, working-age adults, and retirees. Table 2.10 shows the three age groups, the number of people in the town from each age group, and the proportion (%) of people in each age group. Construct a bar graph showing the proportions.

Park City 的人口由儿童、劳动年龄成年人和退休人员组成。表 2.10 显示了三个年龄组、每个年龄组在该镇的人数以及每个年龄组人口所占的比例(%)。请绘制一张显示比例的条形图。

| Age groups | Number of people | Proportion of population |

| 年龄组 | 人数 | 人口比例 |

|--------------------|------------------|--------------------------|

|--------------------|------------------|--------------------------|

| Children | 67,059 | 19% |

| 儿童 | 67,059 | 19% |

| Working-age adults | 152,198 | 43% |

| 劳动年龄成年人 | 152,198 | 43% |

| Retirees | 131,662 | 38% |

| 退休人员 | 131,662 | 38% |

Table 2.10

表 2.10

Problem 问题

The columns in Table 2.11 contain: the race or ethnicity of students in U.S. Public Schools for the class of 2011, percentages for the Advanced Placement examine population for that class, and percentages for the overall student population. Create a bar graph with the student race or ethnicity (qualitative data) on the *x*-axis, and the Advanced Placement examinee population percentages on the *y*-axis.

表 2.11 的各列包含:2011 届美国公立学校学生的种族或族裔、该届先修课程(Advanced Placement)考生群体的百分比,以及全体学生群体的百分比。请制作一张条形图,以学生的种族或族裔(定性数据)为 *x* 轴,以先修课程考生群体的百分比为 *y* 轴。

| Race/Ethnicity | AP Examinee Population | Overall Student Population |

| 种族/族裔 | AP 考生群体 | 全体学生群体 |

|-----------------------------------------------|------------------------|----------------------------|

|-----------------------------------------------|------------------------|----------------------------|

| 1 = Asian, Asian American or Pacific Islander | 10.3% | 5.7% |

| 1 = 亚裔、亚裔美国人或太平洋岛民 | 10.3% | 5.7% |

| 2 = Black or African American | 9.0% | 14.7% |

| 2 = 黑人或非裔美国人 | 9.0% | 14.7% |

| 3 = Hispanic or Latino | 17.0% | 17.6% |

| 3 = 西班牙裔或拉丁裔 | 17.0% | 17.6% |

| 4 = American Indian or Alaska Native | 0.6% | 1.1% |

| 4 = 美洲印第安人或阿拉斯加原住民 | 0.6% | 1.1% |

| 5 = White | 57.1% | 59.2% |

| 5 = 白人 | 57.1% | 59.2% |

| 6 = Not reported/other | 6.0% | 1.7% |

| 6 = 未报告/其他 | 6.0% | 1.7% |

Table 2.11

表 2.11

Solution 解答

Park city is broken down into six voting districts. The table shows the percent of the total registered voter population that lives in each district as well as the percent total of the entire population that lives in each district. Construct a bar graph that shows the registered voter population by district.

帕克城被划分为六个投票区。该表显示了居住在每个选区的登记选民总数占全部登记选民的百分比,以及居住在每个选区的居民占全市居民总数的百分比。请制作一张条形图,按选区显示登记选民群体。

| District | Registered voter population | Overall city population |

| 选区 | 登记选民群体 | 全市居民总体 |

|----------|-----------------------------|-------------------------|

|----------|-----------------------------|-------------------------|

| 1 | 15.5% | 19.4% |

| 1 | 15.5% | 19.4% |

| 2 | 12.2% | 15.6% |

| 2 | 12.2% | 15.6% |

| 3 | 9.8% | 9.0% |

| 3 | 9.8% | 9.0% |

| 4 | 17.4% | 18.5% |

| 4 | 17.4% | 18.5% |

| 5 | 22.8% | 20.7% |

| 5 | 22.8% | 20.7% |

| 6 | 22.3% | 16.8% |

| 6 | 22.3% | 16.8% |

Table 2.12

表 2.12

---

---

2.2 Histograms, Frequency Polygons, and Time Series Graphs 2.2 直方图、频数多边形与时间序列图

For most of the work you do in this book, you will use a histogram to display the data. One advantage of a histogram is that it can readily display large data sets. A rule of thumb is to use a histogram when the data set consists of 100 values or more.

在本书的大部分工作中,你将使用直方图来展示数据。直方图的一个优点是它能够方便地显示大型数据集。经验法则是:当数据集包含 100 个或更多数值时,使用直方图。

A histogram consists of contiguous (adjoining) boxes. It has both a horizontal axis and a vertical axis. The horizontal axis is labeled with what the data represents (for instance, distance from your home to school). The vertical axis is labeled either frequency or relative frequency (or percent frequency or probability). The graph will have the same shape with either label. The histogram (like the stemplot) can give you the shape of the data, the center, and the spread of the data.

直方图由相邻(毗连)的矩形框组成。它既有水平轴,也有垂直轴。水平轴标注数据所代表的内容(例如,从你家到学校的距离)。垂直轴标注频数或相对频数(或百分比频数或概率)。不论使用哪种标注,图形的形状都相同。直方图(与茎叶图类似)能给出数据的形状、中心与离散程度。

The relative frequency is equal to the frequency for an observed value of the data divided by the total number of data values in the sample. (Remember, frequency is defined as the number of times an answer occurs.) If:

相对频数等于某一观测数据值的频数除以样本中数据值的总数。(记住,频数定义为某个答案出现的次数。)若:

then:

则:

$${\text{RF} = \frac{f}{n}}{}$$

$${\text{RF} = \frac{f}{n}}{}$$

For example, if three students in Mr. Ahab's English class of 40 students received from 90% to 100%, then, *f* = 3, *n* = 40, and *RF* = $\frac{f}{n}$ = $\frac{3}{40}$ = 0.075. 7.5% of the students received 90–100%. 90–100% are quantitative measures.

例如,若阿哈布(Ahab)老师所教 40 名学生的英语课上有 3 名学生得分在 90% 到 100% 之间,则 *f* = 3,*n* = 40,*RF* = $\frac{f}{n}$ = $\frac{3}{40}$ = 0.075。7.5% 的学生得分在 90–100%。90–100% 是定量度量。

To construct a histogram, first decide how many bars or intervals, also called classes, represent the data. Many histograms consist of five to 15 bars or classes for clarity. The number of bars needs to be chosen. Choose a starting point for the first interval to be less than the smallest data value. A convenient starting point is a lower value carried out to one more decimal place than the value with the most decimal places. For example, if the value with the most decimal places is 6.1 and this is the smallest value, a convenient starting point is 6.05 (6.1 – 0.05 = 6.05). We say that 6.05 has more precision. If the value with the most decimal places is 2.23 and the lowest value is 1.5, a convenient starting point is 1.495 (1.5 – 0.005 = 1.495). If the value with the most decimal places is 3.234 and the lowest value is 1.0, a convenient starting point is 0.9995 (1.0 – 0.0005 = 0.9995). If all the data happen to be integers and the smallest value is two, then a convenient starting point is 1.5 (2 – 0.5 = 1.5). Also, when the starting point and other boundaries are carried to one additional decimal place, no data value will fall on a boundary. The next two examples go into detail about how to construct a histogram using continuous data and how to create a histogram using discrete data.

制作直方图时,首先要决定用多少个矩形条区间(又称组)来代表数据。为了清晰,许多直方图由 5 到 15 个矩形条或组构成。矩形条的数量需要自行确定。选取第一个区间的起点,使其小于最小的数据值。方便的起点是比小数位数最多的数值再多取一位小数的较低值。例如,若小数位数最多的数值是 6.1 且它是最小值,则方便的起点是 6.05(6.1 – 0.05 = 6.05)。我们说 6.05 具有更高的精度。若小数位数最多的数值是 2.23 而最小值是 1.5,则方便的起点是 1.495(1.5 – 0.005 = 1.495)。若小数位数最多的数值是 3.234 而最小值是 1.0,则方便的起点是 0.9995(1.0 – 0.0005 = 0.9995)。若所有数据恰好都是整数且最小值为 2,则方便的起点是 1.5(2 – 0.5 = 1.5)。此外,当起点及其他边界都多取一位小数时,不会有任何数据值恰好落在边界上。接下来的两个例子将详细说明如何用连续数据制作直方图,以及如何用离散数据制作直方图。

The following data are the heights (in inches to the nearest half inch) of 100 male semiprofessional soccer players. The heights are continuous data, since height is measured.

以下是 100 名男性半职业足球运动员的身高(以英寸计,精确到最近的半英寸)。身高是连续数据,因为身高是测量得到的。

60; 60.5; 61; 61; 61.5

60;60.5;61;61;61.5

63.5; 63.5; 63.5

63.5;63.5;63.5

64; 64; 64; 64; 64; 64; 64; 64.5; 64.5; 64.5; 64.5; 64.5; 64.5; 64.5; 64.5

64;64;64;64;64;64;64;64.5;64.5;64.5;64.5;64.5;64.5;64.5;64.5

66; 66; 66; 66; 66; 66; 66; 66; 66; 66; 66.5; 66.5; 66.5; 66.5; 66.5; 66.5; 66.5; 66.5; 66.5; 66.5; 66.5; 67; 67; 67; 67; 67; 67; 67; 67; 67; 67; 67; 67.5; 67.5; 67.5; 67.5; 67.5; 67.5; 67.5

66;66;66;66;66;66;66;66;66;66;66.5;66.5;66.5;66.5;66.5;66.5;66.5;66.5;66.5;66.5;66.5;67;67;67;67;67;67;67;67;67;67;67;67.5;67.5;67.5;67.5;67.5;67.5;67.5

68; 68; 69; 69; 69; 69; 69; 69; 69; 69; 69; 69; 69.5; 69.5; 69.5; 69.5; 69.5

68;68;69;69;69;69;69;69;69;69;69;69;69.5;69.5;69.5;69.5;69.5

70; 70; 70; 70; 70; 70; 70.5; 70.5; 70.5; 71; 71; 71

70;70;70;70;70;70;70.5;70.5;70.5;71;71;71

72; 72; 72; 72.5; 72.5; 73; 73.5

72;72;72;72.5;72.5;73;73.5

74

74

The smallest data value is 60. Since the data with the most decimal places has one decimal (for instance, 61.5), we want our starting point to have two decimal places. Since the numbers 0.5, 0.05, 0.005, etc. are convenient numbers, use 0.05 and subtract it from 60, the smallest value, for the convenient starting point.

最小的数据值是 60。由于小数位数最多的数据具有一位小数(例如 61.5),我们希望起点具有两位小数。因为 0.5、0.05、0.005 等是方便的数字,取 0.05 并从最小值 60 中减去它,作为方便的起点。

60 – 0.05 = 59.95 which is more precise than, say, 61.5 by one decimal place. The starting point is, then, 59.95.

60 – 0.05 = 59.95,它比例如 61.5 多一位小数,因此更精确。于是起点为 59.95。

The largest value is 74, so 74 + 0.05 = 74.05 is the ending value.

最大值是 74,因此 74 + 0.05 = 74.05 为终点值。

Next, calculate the width of each bar or class interval. To calculate this width, subtract the starting point from the ending value and divide by the number of bars (you must choose the number of bars you desire). Suppose you choose eight bars.

接下来,计算每根矩形条或每个组区间的宽度。要计算该宽度,用终点值减去起点值,再除以矩形条的数量(你必须自行选择所需的矩形条数量)。假设你选择 8 根矩形条。

$$\begin{matrix}{74.05 - 59.95 = 14.1} \\{14.1 \div 8 = 1.76}\end{matrix}$$

$$\begin{matrix}{74.05 - 59.95 = 14.1} \\{14.1 \div 8 = 1.76}\end{matrix}$$

We will round up to two and make each bar or class interval two units wide. Rounding up to two is one way to prevent a value from falling on a boundary. Rounding to the next number is often necessary even if it goes against the standard rules of rounding. For this example, using 1.76 as the width would also work. A guideline that is followed by some for the number of bars or class intervals is to take the square root of the number of data values and then round to the nearest whole number, if necessary. For example, if there are 150 values of data, take the square root of 150 and round to 12 bars or intervals.

我们将向上取整为 2,使每根矩形条或每个组区间的宽度为 2 个单位。向上取整为 2 是一种防止数值落在边界上的方法。即使违背常规的四舍五入规则,取整到下一个数也常常是必要的。在本例中,使用 1.76 作为宽度同样可行。有些人遵循的关于矩形条或组区间数量的准则是:取数据值个数的平方根,必要时再四舍五入到最接近的整数。例如,若有 150 个数据值,则取 150 的平方根并取整为 12 根矩形条或组区间。

The boundaries are:

边界为:

The heights 60 through 61.5 inches are in the interval 59.95–61.95. The heights that are 63.5 are in the interval 61.95–63.95. The heights that are 64 through 64.5 are in the interval 63.95–65.95. The heights 66 through 67.5 are in the interval 65.95–67.95. The heights 68 through 69.5 are in the interval 67.95–69.95. The heights 70 through 71 are in the interval 69.95–71.95. The heights 72 through 73.5 are in the interval 71.95–73.95. The height 74 is in the interval 73.95–75.95.

身高 60 至 61.5 英寸落在 59.95–61.95 区间内。身高为 63.5 的落在 61.95–63.95 区间内。身高 64 至 64.5 的落在 63.95–65.95 区间内。身高 66 至 67.5 的落在 65.95–67.95 区间内。身高 68 至 69.5 的落在 67.95–69.95 区间内。身高 70 至 71 的落在 69.95–71.95 区间内。身高 72 至 73.5 的落在 71.95–73.95 区间内。身高 74 落在 73.95–75.95 区间内。

The following histogram displays the heights on the *x*-axis and relative frequency on the *y*-axis.

下面的直方图以 *x* 轴表示身高,以 *y* 轴表示相对频数。

The following data are the shoe sizes of 50 male students. The sizes are discrete data since shoe size is measured in whole and half units only. Construct a histogram and calculate the width of each bar or class interval. Suppose you choose six bars.

以下是 50 名男学生的鞋号。这些尺码是离散数据,因为鞋号只以整数和半号计量。请制作一张直方图,并计算每根矩形条或每个组区间的宽度。假设你选择 6 根矩形条。

9; 9; 9.5; 9.5; 10; 10; 10; 10; 10; 10; 10.5; 10.5; 10.5; 10.5; 10.5; 10.5; 10.5; 10.5

9;9;9.5;9.5;10;10;10;10;10;10;10.5;10.5;10.5;10.5;10.5;10.5;10.5;10.5

11; 11; 11; 11; 11; 11; 11; 11; 11; 11; 11; 11; 11; 11.5; 11.5; 11.5; 11.5; 11.5; 11.5; 11.5

11;11;11;11;11;11;11;11;11;11;11;11;11;11.5;11.5;11.5;11.5;11.5;11.5;11.5

12; 12; 12; 12; 12; 12; 12; 12.5; 12.5; 12.5; 12.5; 14

12;12;12;12;12;12;12;12.5;12.5;12.5;12.5;14

Create a histogram for the following data: the number of books bought by 50 part-time college students at ABC College. The number of books is discrete data, since books are counted.

根据以下数据制作直方图:ABC 学院 50 名兼职大学生购买的书籍数量。书籍数量是离散数据,因为书是按个数计的。

1; 1; 1; 1; 1; 1; 1; 1; 1; 1; 1

1;1;1;1;1;1;1;1;1;1;1

2; 2; 2; 2; 2; 2; 2; 2; 2; 2

2;2;2;2;2;2;2;2;2;2

3; 3; 3; 3; 3; 3; 3; 3; 3; 3; 3; 3; 3; 3; 3; 3

3;3;3;3;3;3;3;3;3;3;3;3;3;3;3;3

4; 4; 4; 4; 4; 4

4;4;4;4;4;4

5; 5; 5; 5; 5

5;5;5;5;5

6; 6

6;6

Eleven students buy one book. Ten students buy two books. Sixteen students buy three books. Six students buy four books. Five students buy five books. Two students buy six books.

11 名学生买 1 本书,10 名学生买 2 本书,16 名学生买 3 本书,6 名学生买 4 本书,5 名学生买 5 本书,2 名学生买 6 本书。

Because the data are integers, subtract 0.5 from 1, the smallest data value and add 0.5 to 6, the largest data value. Then the starting point is 0.5 and the ending value is 6.5.

由于数据是整数,从最小数据值 1 减去 0.5,并向最大数据值 6 加上 0.5。于是起点为 0.5,终点值为 6.5。

Problem 问题

Next, calculate the width of each bar or class interval. If the data are discrete and there are not too many different values, a width that places the data values in the middle of the bar or class interval is the most convenient. Since the data consist of the numbers 1, 2, 3, 4, 5, 6, and the starting point is 0.5, a width of one places the 1 in the middle of the interval from 0.5 to 1.5, the 2 in the middle of the interval from 1.5 to 2.5, the 3 in the middle of the interval from 2.5 to 3.5, the 4 in the middle of the interval from \_\_\_\_\_\_\_ to \_\_\_\_\_\_\_, the 5 in the middle of the interval from \_\_\_\_\_\_\_ to \_\_\_\_\_\_\_, and the \_\_\_\_\_\_\_ in the middle of the interval from \_\_\_\_\_\_\_ to \_\_\_\_\_\_\_ .

接下来,计算每根矩形条或每个组区间的宽度。若数据是离散的,且不同取值不太多,则使数据值落在矩形条或组区间中部的宽度最为方便。由于数据由数字 1、2、3、4、5、6 组成,且起点为 0.5,宽度为 1 时,1 落在 0.5 到 1.5 区间的中部,2 落在 1.5 到 2.5 区间的中部,3 落在 2.5 到 3.5 区间的中部,4 落在 \_\_\_\_\_\_\_ 到 \_\_\_\_\_\_\_ 区间的中部,5 落在 \_\_\_\_\_\_\_ 到 \_\_\_\_\_\_\_ 区间的中部,\_\_\_\_\_\_\_ 落在 \_\_\_\_\_\_\_ 到 \_\_\_\_\_\_\_ 区间的中部。

Solution 解答

Calculate the number of bars as follows:

按如下方式计算矩形条的数量:

$$\begin{matrix}{6.5 - 0.5 = 6} \\{6 \div 1 = 6}\end{matrix}$$

$$\begin{matrix}{6.5 - 0.5 = 6} \\{6 \div 1 = 6}\end{matrix}$$

where 1 is the width of a bar. Therefore, bars = 6.

其中 1 是一根矩形条的宽度。因此,矩形条数 = 6。

The following histogram displays the number of books on the *x*-axis and the frequency on the *y*-axis.

下面的直方图以 *x* 轴表示书籍数量,以 *y* 轴表示频数。

Go to Appendix G Notes for the TI-83, 83+, 84, 84+ Calculators. There are calculator instructions for entering data and for creating a customized histogram. Create the histogram for Example 2.8.

参见附录 G 中关于 TI-83、83+、84、84+ 计算器的说明。其中有输入数据以及制作自定义直方图的 calculator 操作说明。请制作例 2.8 的直方图。

The following data are the number of sports played by 50 student athletes. The number of sports is discrete data since sports are counted.

以下是 50 名学生运动员所参加的运动项目数。运动项目数是离散数据,因为运动项目是按个数计的。

1; 1; 1; 1; 1; 1; 1; 1; 1; 1; 1; 1; 1; 1; 1; 1; 1; 1; 1; 1

1;1;1;1;1;1;1;1;1;1;1;1;1;1;1;1;1;1;1;1

2; 2; 2; 2; 2; 2; 2; 2; 2; 2; 2; 2; 2; 2; 2; 2; 2; 2; 2; 2; 2; 2

2;2;2;2;2;2;2;2;2;2;2;2;2;2;2;2;2;2;2;2;2;2

3; 3; 3; 3; 3; 3; 3; 3

3;3;3;3;3;3;3;3

20 student athletes play one sport. 22 student athletes play two sports. Eight student athletes play three sports.

20 名学生运动员参加 1 项运动,22 名参加 2 项运动,8 名参加 3 项运动。

*Fill in the blanks for the following sentence.* Since the data consist of the numbers 1, 2, 3, and the starting point is 0.5, a width of one places the 1 in the middle of the interval 0.5 to \_\_\_\_\_, the 2 in the middle of the interval from \_\_\_\_\_ to \_\_\_\_\_, and the 3 in the middle of the interval from \_\_\_\_\_ to \_\_\_\_\_.

*补全下面句子的空白处。* 由于数据由数字 1、2、3 组成,且起点为 0.5,宽度为 1 时,1 落在 0.5 到 \_\_\_\_\_ 区间的中部,2 落在 \_\_\_\_\_ 到 \_\_\_\_\_ 区间的中部,3 落在 \_\_\_\_\_ 到 \_\_\_\_\_ 区间的中部。

Problem 问题

Using this data set, construct a histogram.

使用该数据集制作一张直方图。

| Number of Hours My Classmates Spent Playing Video Games on Weekends | | | | |

| 我同学周末玩电子游戏所花小时数 | | | | |

|---------------------------------------------------------------------|------|------|-------|-------|

|---------------------------------------------------------------------|------|------|-------|-------|

| 9.95 | 10 | 2.25 | 16.75 | 0 |

| 9.95 | 10 | 2.25 | 16.75 | 0 |

| 19.5 | 22.5 | 7.5 | 15 | 12.75 |

| 19.5 | 22.5 | 7.5 | 15 | 12.75 |

| 5.5 | 11 | 10 | 20.75 | 17.5 |

| 5.5 | 11 | 10 | 20.75 | 17.5 |

| 23 | 21.9 | 24 | 23.75 | 18 |

| 23 | 21.9 | 24 | 23.75 | 18 |

| 20 | 15 | 22.9 | 18.8 | 20.5 |

| 20 | 15 | 22.9 | 18.8 | 20.5 |

Table 2.13

表 2.13

Solution 解答

Some values in this data set fall on boundaries for the class intervals. A value is counted in a class interval if it falls on the left boundary, but not if it falls on the right boundary. Different researchers may set up histograms for the same data in different ways. There is more than one correct way to set up a histogram.

该数据集中的某些数值恰好落在组区间的边界上。若一个数值落在左边界上,则计入该组区间;若落在右边界上,则不计入。不同的研究者可能以不同方式为同一数据建立直方图。建立直方图的正确方式不止一种。

The following data represent the number of employees at various restaurants in New York City. Using this data, create a histogram.

以下数据表示纽约市各家餐馆的员工人数。请使用该数据制作一张直方图。

22; 35; 15; 26; 40; 28; 18; 20; 25; 34; 39; 42; 24; 22; 19; 27; 22; 34; 40; 20; 38; and 28

22;35;15;26;40;28;18;20;25;34;39;42;24;22;19;27;22;34;40;20;38;以及 28

Use 10–19 as the first interval.

以 10–19 作为第一个区间。

Count the money (bills and change) in your pocket or purse. Your instructor will record the amounts. As a class, construct a histogram displaying the data. Discuss how many intervals you think is appropriate. You may want to experiment with the number of intervals.

数一数你口袋或钱包里的钱(纸币和硬币)。你的老师会记录这些金额。以班级为单位,制作一张直方图来展示这些数据。讨论你认为多少个区间合适。你可以尝试不同的区间数量。

Frequency Polygons 频数多边形

Frequency polygons are analogous to line graphs, and just as line graphs make continuous data visually easy to interpret, so too do frequency polygons.

频数多边形类似于线图;正如线图使连续数据在视觉上易于解读,频数多边形也是如此。

To construct a frequency polygon, first examine the data and decide on the number of intervals, or class intervals, to use on the *x*-axis and *y*-axis. After choosing the appropriate ranges, begin plotting the data points. After all the points are plotted, draw line segments to connect them.

要构建频数多边形,首先考察数据并决定在 *x* 轴和 *y* 轴上使用的区间数(即组距)。选定合适的范围后,开始描点。所有点描完后,用线段将它们连接起来。

A frequency polygon was constructed from the frequency table below.

下面的频数多边形是根据该频数表构建的。

| Frequency Distribution for Calculus Final Test Scores | | | |

| 微积分期末考试成绩频数分布 | | | |

|-------------------------------------------------------|-------------|-----------|----------------------|

|-------------------------------------------------------|-------------|-----------|----------------------|

| Lower Bound | Upper Bound | Frequency | Cumulative Frequency |

| 下界 | 上界 | 频数 | 累积频数 |

| 49.5 | 59.5 | 5 | 5 |

| 49.5 | 59.5 | 5 | 5 |

| 59.5 | 69.5 | 10 | 15 |

| 59.5 | 69.5 | 10 | 15 |

| 69.5 | 79.5 | 30 | 45 |

| 69.5 | 79.5 | 30 | 45 |

| 79.5 | 89.5 | 40 | 85 |

| 79.5 | 89.5 | 40 | 85 |

| 89.5 | 99.5 | 15 | 100 |

| 89.5 | 99.5 | 15 | 100 |

Table 2.14

表 2.14

The first label on the *x*-axis is 44.5. This represents an interval extending from 39.5 to 49.5. Since the lowest test score is 54.5, this interval is used only to allow the graph to touch the *x*-axis. The point labeled 54.5 represents the next interval, or the first “real” interval from the table, and contains five scores. This reasoning is followed for each of the remaining intervals with the point 104.5 representing the interval from 99.5 to 109.5. Again, this interval contains no data and is only used so that the graph will touch the *x*-axis. Looking at the graph, we say that this distribution is skewed because one side of the graph does not mirror the other side.

*x* 轴上的第一个标记是 44.5,它代表一个从 39.5 延伸到 49.5 的区间。由于最低考试分数为 54.5,该区间仅用于让图形触及 *x* 轴。标记为 54.5 的点代表下一个区间,即表中第一个"真实"区间,包含 5 个分数。对其余各区间依此类推,其中点 104.5 代表从 99.5 到 109.5 的区间。同样,该区间不含数据,仅用于使图形触及 *x* 轴。观察该图形,我们说该分布是偏斜的,因为图形的一侧并不与另一侧对称。

Construct a frequency polygon of U.S. Presidents’ ages at inauguration shown in Table 2.15.

根据表 2.15 中美国总统就职时的年龄,构建频数多边形。

| Age at Inauguration | Frequency |

| 就职年龄 | 频数 |

|---------------------|-----------|

|---------------------|-----------|

| 41.5–46.5 | 4 |

| 41.5–46.5 | 4 |

| 46.5–51.5 | 11 |

| 46.5–51.5 | 11 |

| 51.5–56.5 | 14 |

| 51.5–56.5 | 14 |

| 56.5–61.5 | 9 |

| 56.5–61.5 | 9 |

| 61.5–66.5 | 4 |

| 61.5–66.5 | 4 |

| 66.5–71.5 | 2 |

| 66.5–71.5 | 2 |

Table 2.15

表 2.15

Frequency polygons are useful for comparing distributions. This is achieved by overlaying the frequency polygons drawn for different data sets.

频数多边形可用于比较分布。方法是把针对不同的数据集绘制的频数多边形叠加在一起。

We will construct an overlay frequency polygon comparing the scores from Example 2.10 with the students’ final numeric grade.

我们将构建一个叠加频数多边形,把例 2.10 的分数与学生的最终数字成绩进行比较。

| Frequency Distribution for Calculus Final Test Scores | | | |

| 微积分期末考试成绩频数分布 | | | |

|-------------------------------------------------------|-------------|-----------|----------------------|

|-------------------------------------------------------|-------------|-----------|----------------------|

| Lower Bound | Upper Bound | Frequency | Cumulative Frequency |

| 下界 | 上界 | 频数 | 累积频数 |

| 49.5 | 59.5 | 5 | 5 |

| 49.5 | 59.5 | 5 | 5 |

| 59.5 | 69.5 | 10 | 15 |

| 59.5 | 69.5 | 10 | 15 |

| 69.5 | 79.5 | 30 | 45 |

| 69.5 | 79.5 | 30 | 45 |

| 79.5 | 89.5 | 40 | 85 |

| 79.5 | 89.5 | 40 | 85 |

| 89.5 | 99.5 | 15 | 100 |

| 89.5 | 99.5 | 15 | 100 |

Table 2.16

表 2.16

| Frequency Distribution for Calculus Final Grades | | | |

| 微积分期末成绩频数分布 | | | |

|--------------------------------------------------|-------------|-----------|----------------------|

|--------------------------------------------------|-------------|-----------|----------------------|

| Lower Bound | Upper Bound | Frequency | Cumulative Frequency |

| 下界 | 上界 | 频数 | 累积频数 |

| 49.5 | 59.5 | 10 | 10 |

| 49.5 | 59.5 | 10 | 10 |

| 59.5 | 69.5 | 10 | 20 |

| 59.5 | 69.5 | 10 | 20 |

| 69.5 | 79.5 | 30 | 50 |

| 69.5 | 79.5 | 30 | 50 |

| 79.5 | 89.5 | 45 | 95 |

| 79.5 | 89.5 | 45 | 95 |

| 89.5 | 99.5 | 5 | 100 |

| 89.5 | 99.5 | 5 | 100 |

Table 2.17

表 2.17

Suppose that we want to study the temperature range of a region for an entire month. Every day at noon we note the temperature and write this down in a log. A variety of statistical studies could be done with this data. We could find the mean or the median temperature for the month. We could construct a histogram displaying the number of days that temperatures reach a certain range of values. However, all of these methods ignore a portion of the data that we have collected.

假设我们想研究某个地区一整月的温度范围。每天中午我们记录温度并写入日志。利用这些数据可以做多种统计研究:可以求当月温度的均值或中位数,也可以构建直方图来显示温度达到某个数值范围的天数。然而,所有这些方法都忽略了我们所收集数据的一部分。

One feature of the data that we may want to consider is that of time. Since each date is paired with the temperature reading for the day, we don‘t have to think of the data as being random. We can instead use the times given to impose a chronological order on the data. A graph that recognizes this ordering and displays the changing temperature as the month progresses is called a time series graph.

我们可能想考虑的数据的一个特征是时间。由于每个日期都与当天的温度读数配对,我们不必把数据看作随机的。相反,我们可以利用给定的时间给数据强加一个时间顺序。一种识别这种顺序并在月份推进过程中显示温度变化情况的图形,称为时间序列图。

Constructing a Time Series Graph 构建时间序列图

To construct a time series graph, we must look at both pieces of our paired data set. We start with a standard Cartesian coordinate system. The horizontal axis is used to plot the date or time increments, and the vertical axis is used to plot the values of the variable that we are measuring. By doing this, we make each point on the graph correspond to a date and a measured quantity. The points on the graph are typically connected by straight lines in the order in which they occur.

要构建时间序列图,必须考察配对数据集的两个部分。我们先采用标准的笛卡尔坐标系:横轴用来标绘日期或时间增量,纵轴用来标绘所测变量的值。这样,图形上的每个点都对应一个日期和一个被测数量。图上的点通常按发生的先后顺序用直线连接起来。

Problem 问题

The following data shows the Annual Consumer Price Index, each month, for ten years. Construct a time series graph for the Annual Consumer Price Index data only.

以下数据给出了十年间各月的年度消费者价格指数。仅为该年度消费者价格指数数据构建时间序列图。

| Year | Jan | Feb | Mar | Apr | May | Jun | Jul |

| 年份 | 1月 | 2月 | 3月 | 4月 | 5月 | 6月 | 7月 |

|----------|---------|---------|---------|---------|---------|---------|---------|

|----------|---------|---------|---------|---------|---------|---------|---------|

| 2003 | 181.7 | 183.1 | 184.2 | 183.8 | 183.5 | 183.7 | 183.9 |

| 2003 | 181.7 | 183.1 | 184.2 | 183.8 | 183.5 | 183.7 | 183.9 |

| 2004 | 185.2 | 186.2 | 187.4 | 188.0 | 189.1 | 189.7 | 189.4 |

| 2004 | 185.2 | 186.2 | 187.4 | 188.0 | 189.1 | 189.7 | 189.4 |

| 2005 | 190.7 | 191.8 | 193.3 | 194.6 | 194.4 | 194.5 | 195.4 |

| 2005 | 190.7 | 191.8 | 193.3 | 194.6 | 194.4 | 194.5 | 195.4 |

| 2006 | 198.3 | 198.7 | 199.8 | 201.5 | 202.5 | 202.9 | 203.5 |

| 2006 | 198.3 | 198.7 | 199.8 | 201.5 | 202.5 | 202.9 | 203.5 |

| 2007 | 202.416 | 203.499 | 205.352 | 206.686 | 207.949 | 208.352 | 208.299 |

| 2007 | 202.416 | 203.499 | 205.352 | 206.686 | 207.949 | 208.352 | 208.299 |

| 2008 | 211.080 | 211.693 | 213.528 | 214.823 | 216.632 | 218.815 | 219.964 |

| 2008 | 211.080 | 211.693 | 213.528 | 214.823 | 216.632 | 218.815 | 219.964 |

| 2009 | 211.143 | 212.193 | 212.709 | 213.240 | 213.856 | 215.693 | 215.351 |

| 2009 | 211.143 | 212.193 | 212.709 | 213.240 | 213.856 | 215.693 | 215.351 |

| 2010 | 216.687 | 216.741 | 217.631 | 218.009 | 218.178 | 217.965 | 218.011 |

| 2010 | 216.687 | 216.741 | 217.631 | 218.009 | 218.178 | 217.965 | 218.011 |

| 2011 | 220.223 | 221.309 | 223.467 | 224.906 | 225.964 | 225.722 | 225.922 |

| 2011 | 220.223 | 221.309 | 223.467 | 224.906 | 225.964 | 225.722 | 225.922 |

| 2012 | 226.665 | 227.663 | 229.392 | 230.085 | 229.815 | 229.478 | 229.104 |

| 2012 | 226.665 | 227.663 | 229.392 | 230.085 | 229.815 | 229.478 | 229.104 |

Table 2.18

表 2.18

| Year | Aug | Sep | Oct | Nov | Dec | Annual |

| 年份 | 8月 | 9月 | 10月 | 11月 | 12月 | 年度 |

|----------|---------|---------|---------|---------|---------|---------|

|----------|---------|---------|---------|---------|---------|---------|

| 2003 | 184.6 | 185.2 | 185.0 | 184.5 | 184.3 | 184.0 |

| 2003 | 184.6 | 185.2 | 185.0 | 184.5 | 184.3 | 184.0 |

| 2004 | 189.5 | 189.9 | 190.9 | 191.0 | 190.3 | 188.9 |

| 2004 | 189.5 | 189.9 | 190.9 | 191.0 | 190.3 | 188.9 |

| 2005 | 196.4 | 198.8 | 199.2 | 197.6 | 196.8 | 195.3 |

| 2005 | 196.4 | 198.8 | 199.2 | 197.6 | 196.8 | 195.3 |

| 2006 | 203.9 | 202.9 | 201.8 | 201.5 | 201.8 | 201.6 |

| 2006 | 203.9 | 202.9 | 201.8 | 201.5 | 201.8 | 201.6 |

| 2007 | 207.917 | 208.490 | 208.936 | 210.177 | 210.036 | 207.342 |

| 2007 | 207.917 | 208.490 | 208.936 | 210.177 | 210.036 | 207.342 |

| 2008 | 219.086 | 218.783 | 216.573 | 212.425 | 210.228 | 215.303 |

| 2008 | 219.086 | 218.783 | 216.573 | 212.425 | 210.228 | 215.303 |

| 2009 | 215.834 | 215.969 | 216.177 | 216.330 | 215.949 | 214.537 |

| 2009 | 215.834 | 215.969 | 216.177 | 216.330 | 215.949 | 214.537 |

| 2010 | 218.312 | 218.439 | 218.711 | 218.803 | 219.179 | 218.056 |

| 2010 | 218.312 | 218.439 | 218.711 | 218.803 | 219.179 | 218.056 |

| 2011 | 226.545 | 226.889 | 226.421 | 226.230 | 225.672 | 224.939 |

| 2011 | 226.545 | 226.889 | 226.421 | 226.230 | 225.672 | 224.939 |

| 2012 | 230.379 | 231.407 | 231.317 | 230.221 | 229.601 | 229.594 |

| 2012 | 230.379 | 231.407 | 231.317 | 230.221 | 229.601 | 229.594 |

Table 2.19

表 2.19

Solution 解答

The following table is a portion of a data set from www.worldbank.org. Use the table to construct a time series graph for CO2 emissions for the United States.

下表是来自 www.worldbank.org 的数据集的一部分。利用该表为美国的 CO2 排放构建时间序列图。

| CO2 Emissions | | | |

| CO₂ 排放 | | | |

|---------------|---------|----------------|---------------|

|---------------|---------|----------------|---------------|

| | Ukraine | United Kingdom | United States |

| | 乌克兰 | 英国 | 美国 |

| 2003 | 352,259 | 540,640 | 5,681,664 |

| 2003 | 352,259 | 540,640 | 5,681,664 |

| 2004 | 343,121 | 540,409 | 5,790,761 |

| 2004 | 343,121 | 540,409 | 5,790,761 |

| 2005 | 339,029 | 541,990 | 5,826,394 |

| 2005 | 339,029 | 541,990 | 5,826,394 |

| 2006 | 327,797 | 542,045 | 5,737,615 |

| 2006 | 327,797 | 542,045 | 5,737,615 |

| 2007 | 328,357 | 528,631 | 5,828,697 |

| 2007 | 328,357 | 528,631 | 5,828,697 |

| 2008 | 323,657 | 522,247 | 5,656,839 |

| 2008 | 323,657 | 522,247 | 5,656,839 |

| 2009 | 272,176 | 474,579 | 5,299,563 |

| 2009 | 272,176 | 474,579 | 5,299,563 |

Table 2.20

表 2.20

Uses of a Time Series Graph 时间序列图的用途

Time series graphs are important tools in various applications of statistics. When recording values of the same variable over an extended period of time, sometimes it is difficult to discern any trend or pattern. However, once the same data points are displayed graphically, some features jump out. Time series graphs make trends easy to spot.

时间序列图是统计各种应用中的重要工具。当在较长时期内记录同一变量的取值时,有时很难辨出任何趋势或规律。然而,一旦把这些相同的数据点用图形展示出来,某些特征就会凸显。时间序列图使趋势易于发现。

---

---

2.3 Measures of the Location of the Data 2.3 数据的位置度量

The common measures of location are quartiles and percentiles

常见的位置度量是四分位数和百分位数。

Quartiles are special percentiles. The first quartile, *Q*1, is the same as the 25th percentile, and the third quartile, *Q*3, is the same as the 75th percentile. The median, *M*, is called both the second quartile and the 50th percentile.

四分位数是特殊的百分位数。第一四分位数 *Q*1 等同于第 25 百分位数,第三四分位数 *Q*3 等同于第 75 百分位数。中位数 *M* 既称为第二四分位数,也称为第 50 百分位数。

To calculate quartiles and percentiles, the data must be ordered from smallest to largest. Quartiles divide ordered data into quarters. Percentiles divide ordered data into hundredths. To score in the 90th percentile of an exam does not mean, necessarily, that you received 90% on a test. It means that 90% of test scores are the same or less than your score and 10% of the test scores are the same or greater than your test score.

要计算四分位数和百分位数,数据必须按从小到大的顺序排列。四分位数把有序数据分成四份,百分位数把有序数据分成一百份。在一次考试中处于第 90 百分位数,并不一定意味着你考了 90 分。它的含义是:90% 的考试分数低于或等于你的分数,而 10% 的考试分数高于或等于你的分数。

Percentiles are useful for comparing values. For this reason, universities and colleges use percentiles extensively. One instance in which colleges and universities use percentiles is when SAT results are used to determine a minimum testing score that will be used as an acceptance factor. For example, suppose Duke accepts SAT scores at or above the 75th percentile. That translates into a score of at least 1220.

百分位数便于比较数值,因此大学和学院广泛使用百分位数。高校使用百分位数的一个例子,是用 SAT 成绩来确定作为录取依据的最低分数。例如,假设杜克大学接受第 75 百分位数及以上(含)的 SAT 成绩,这相当于至少 1220 分。

Percentiles are mostly used with very large populations. Therefore, if you were to say that 90% of the test scores are less (and not the same or less) than your score, it would be acceptable because removing one particular data value is not significant.

百分数位数主要用于非常大的总体。因此,如果你说 90% 的考试分数低于(而非低于或等于)你的分数,也是可以接受的,因为去掉某一个特定的数据值影响不大。

The median is a number that measures the "center" of the data. You can think of the median as the "middle value," but it does not actually have to be one of the observed values. It is a number that separates ordered data into halves. Half the values are the same number or smaller than the median, and half the values are the same number or larger. For example, consider the following data.

中位数是一个度量数据"中心"的数值。你可以把中位数看作"中间值",但它实际上不必是某个观测值。它是一个把有序数据分成两半的数值:一半的值小于或等于中位数,另一半的值大于或等于中位数。例如,考虑以下数据。

1; 11.5; 6; 7.2; 4; 8; 9; 10; 6.8; 8.3; 2; 2; 10; 1

1;11.5;6;7.2;4;8;9;10;6.8;8.3;2;2;10;1

Ordered from smallest to largest:

从小到大排列:

1; 1; 2; 2; 4; 6; 6.8; 7.2; 8; 8.3; 9; 10; 10; 11.5

1;1;2;2;4;6;6.8;7.2;8;8.3;9;10;10;11.5

Since there are 14 observations, the median is between the seventh value, 6.8, and the eighth value, 7.2. To find the median, add the two values together and divide by two.

由于共有 14 个观测值,中位数位于第 7 个值 6.8 与第 8 个值 7.2 之间。求中位数时,把这两个值相加再除以 2。

$$\begin{matrix}{6.8 + 7.2 = 14} \\{14 \div 2 = 7}\end{matrix}$$

$$\begin{matrix}{6.8 + 7.2 = 14} \\{14 \div 2 = 7}\end{matrix}$$

The median is seven. Half of the values are smaller than seven and half of the values are larger than seven.

中位数为 7。一半的值小于 7,另一半的值大于 7。

Quartiles are numbers that separate the data into quarters. Quartiles may or may not be part of the data. To find the quartiles, first find the median or second quartile. The first quartile, *Q*1, is the middle value of the lower half of the data, and the third quartile, *Q*3, is the middle value, or median, of the upper half of the data. To get the idea, consider the same data set:

四分位数是把数据分成四份的数值。四分位数可能是、也可能不是数据本身的一部分。求四分位数时,先求中位数(第二四分位数)。第一四分位数 *Q*1 是数据下半部分的中间值,第三四分位数 *Q*3 是数据上半部分的中间值(即中位数)。为理解其含义,考虑同一数据集:

1; 1; 2; 2; 4; 6; 6.8; 7.2; 8; 8.3; 9; 10; 10; 11.5

1;1;2;2;4;6;6.8;7.2;8;8.3;9;10;10;11.5

The median or second quartile is seven. The lower half of the data are 1, 1, 2, 2, 4, 6, 6.8. The middle value of the lower half is two.

中位数(即第二四分位数)为 7。数据的下半部分是 1、1、2、2、4、6、6.8,其下半部分的中间值是 2。

1; 1; 2; 2; 4; 6; 6.8

1;1;2;2;4;6;6.8

The number two, which is part of the data, is the first quartile. One-fourth of the entire sets of values are the same as or less than two and three-fourths of the values are more than two.

数值 2 是数据的一部分,它是第一四分位数。整个数据集的四分之一的值小于或等于 2,四分之三的值大于 2。

The upper half of the data is 7.2, 8, 8.3, 9, 10, 10, 11.5. The middle value of the upper half is nine.

数据的上半部分是 7.2、8、8.3、9、10、10、11.5,其上半部分的中间值是 9。

The third quartile, *Q*3, is nine. Three-fourths (75%) of the ordered data set are less than nine. One-fourth (25%) of the ordered data set are greater than nine. The third quartile is part of the data set in this example.

第三四分位数 *Q*3 为 9。有序数据集的四分之三(75%)小于 9,四分之一(25%)大于 9。在本例中,第三四分位数是该数据集的一部分。

The interquartile range is a number that indicates the spread of the middle half or the middle 50% of the data. It is the difference between the third quartile (*Q*3) and the first quartile (*Q*1).

四分位距(IQR)是一个指示数据中间一半(即中间 50%)离散程度的数值,等于第三四分位数 (*Q*3) 与第一四分位数 (*Q*1) 之差。

*IQR* = *Q*3 – *Q*1

*IQR* = *Q*3 – *Q*1

The *IQR* can help to determine potential outliers. **A value is suspected to be a potential outlier if it is less than (1.5)(*IQR*) below the first quartile or more than (1.5)(*IQR*) above the third quartile**. Potential outliers always require further investigation.

*IQR* 有助于判断潜在的离群值。**若某个值小于第一四分位数以下 (1.5)(*IQR*),或大于第三四分位数以上 (1.5)(*IQR*),则该值被怀疑为潜在离群值**。潜在离群值总是需要进一步调查。

A potential outlier is a data point that is significantly different from the other data points. These special data points may be errors or some kind of abnormality or they may be a key to understanding the data.

潜在离群值是与其他数据点明显不同的数据点。这些特殊数据点可能是错误,也可能是某种异常,也可能是理解数据的关键。

Problem 问题

For the following 13 real estate prices, calculate the *IQR* and determine if any prices are potential outliers. Prices are in dollars.

对以下 13 个房地产价格,计算 *IQR* 并判断是否有价格属于潜在离群值。价格单位为美元。

389,950; 230,500; 158,000; 479,000; 639,000; 114,950; 5,500,000; 387,000; 659,000; 529,000; 575,000; 488,800; 1,095,000

389,950;230,500;158,000;479,000;639,000;114,950;5,500,000;387,000;659,000;529,000;575,000;488,800;1,095,000

Solution 解答

Order the data from smallest to largest.

将数据从小到大排列。

114,950; 158,000; 230,500; 387,000; 389,950; 479,000; 488,800; 529,000; 575,000; 639,000; 659,000; 1,095,000; 5,500,000

114,950;158,000;230,500;387,000;389,950;479,000;488,800;529,000;575,000;639,000;659,000;1,095,000;5,500,000

*M* = 488,800

*M* = 488,800

*Q*1 = $\frac{\text{230,500~+~387,000}}{2}$ = 308,750

*Q*1 = $\frac{\text{230,500~+~387,000}}{2}$ = 308,750

*Q*3 = $\frac{\text{639,000~+~659,000}}{2}$ = 649,000

*Q*3 = $\frac{\text{639,000~+~659,000}}{2}$ = 649,000

*IQR* = 649,000 – 308,750 = 340,250

*IQR* = 649,000 – 308,750 = 340,250

(1.5)(*IQR*) = (1.5)(340,250) = 510,375

(1.5)(*IQR*) = (1.5)(340,250) = 510,375

*Q*1 – (1.5)(*IQR*) = 308,750 – 510,375 = –201,625

*Q*1 – (1.5)(*IQR*) = 308,750 – 510,375 = –201,625

*Q*3 + (1.5)(*IQR*) = 649,000 + 510,375 = 1,159,375

*Q*3 + (1.5)(*IQR*) = 649,000 + 510,375 = 1,159,375

No house price is less than –201,625. However, 5,500,000 is more than 1,159,375. Therefore, 5,500,000 is a potential outlier.

没有房价小于 –201,625。但 5,500,000 大于 1,159,375,因此 5,500,000 是潜在离群值。

For the following 11 salaries, calculate the *IQR* and determine if any salaries are outliers. The salaries are in dollars.

对以下 11 个薪水,计算 *IQR* 并判断是否有薪水属于离群值。薪水单位为美元。

\$33,000; \$64,500; \$28,000; \$54,000; \$72,000; \$68,500; \$69,000; \$42,000; \$54,000; \$120,000; \$40,500

\$33,000;\$64,500;\$28,000;\$54,000;\$72,000;\$68,500;\$69,000;\$42,000;\$54,000;\$120,000;\$40,500

Problem 问题

For the two data sets in the test scores example, find the following:

针对考试成绩例子中的两个数据集,求下列内容:

1. The interquartile range. Compare the two interquartile ranges.

1. 四分位距。比较两个四分位距。

2. Any outliers in either set.

2. 任一集合中的离群值。

Solution 解答

The five number summary for the day and night classes is

日间班和夜间班的五数概括如下

| | Minimum | *Q*1 | Median | *Q*3 | Maximum |

| | 最小值 | *Q*1 | 中位数 | *Q*3 | 最大值 |

|-----------|---------|-----------------|--------|-----------------|---------|

|-----------|---------|-----------------|--------|-----------------|---------|

| Day | 32 | 56 | 74.5 | 82.5 | 99 |

| 日间 | 32 | 56 | 74.5 | 82.5 | 99 |

| Night | 25.5 | 78 | 81 | 89 | 98 |

| 夜间 | 25.5 | 78 | 81 | 89 | 98 |

Table 2.21

表 2.21

1. The IQR for the day group is *Q*3 – *Q*1 = 82.5 – 56 = 26.5

1. 日间组的四分位距(IQR)为 *Q*3 – *Q*1 = 82.5 – 56 = 26.5

The IQR for the night group is *Q*3 – *Q*1 = 89 – 78 = 11

夜间组的四分位距(IQR)为 *Q*3 – *Q*1 = 89 – 78 = 11

The interquartile range (the spread or variability) for the day class is larger than the night class *IQR*. This suggests more variation will be found in the day class’s class test scores.

日间班数据的四分位距(离散程度或变异性)大于夜间班的 IQR。这表明日间班的测验成绩会有更大的差异。

2. Day class outliers are found using the IQR times 1.5 rule. So,

2. 日间班的离群值用 IQR 乘以 1.5 的准则来寻找。即,

Since the minimum and maximum values for the day class are greater than 16.25 and less than 122.25, there are no outliers.

由于日间班的最小值和最大值都大于 16.25 且小于 122.25,因此没有离群值。

Night class outliers are calculated as:

夜间班的离群值计算如下:

For this class, any test score less than 61.5 is an outlier. Therefore, the scores of 45 and 25.5 are outliers. Since no test score is greater than 105.5, there is no upper end outlier.

对夜间班而言,任何低于 61.5 的测验成绩都是离群值。因此,45 分和 25.5 分是离群值。由于没有测验成绩高于 105.5,所以不存在高端离群值。

Find the interquartile range for the following two data sets and compare them.

求下列两个数据集的四分位距并作比较。

Test Scores for Class *A*

A 班测试成绩

69; 96; 81; 79; 65; 76; 83; 99; 89; 67; 90; 77; 85; 98; 66; 91; 77; 69; 80; 94

69;96;81;79;65;76;83;99;89;67;90;77;85;98;66;91;77;69;80;94

Test Scores for Class *B*

B 班测试成绩

90; 72; 80; 92; 90; 97; 92; 75; 79; 68; 70; 80; 99; 95; 78; 73; 71; 68; 95; 100

90;72;80;92;90;97;92;75;79;68;70;80;99;95;78;73;71;68;95;100

Fifty statistics students were asked how much sleep they get per school night (rounded to the nearest hour). The results were:

50 名统计学专业的学生被问及每个校夜睡多少小时(四舍五入到最近的整数)。结果如下:

| AMOUNT OF SLEEP PER SCHOOL NIGHT (HOURS) | FREQUENCY | RELATIVE FREQUENCY | CUMULATIVE RELATIVE FREQUENCY |

| 每校夜睡眠时长(小时) | 频数 | 相对频数 | 累积相对频数 |

|------------------------------------------|-----------|--------------------|-------------------------------|

|------------------------------------------|-----------|--------------------|-------------------------------|

| 4 | 2 | 0.04 | 0.04 |

| 4 | 2 | 0.04 | 0.04 |

| 5 | 5 | 0.10 | 0.14 |

| 5 | 5 | 0.10 | 0.14 |

| 6 | 7 | 0.14 | 0.28 |

| 6 | 7 | 0.14 | 0.28 |

| 7 | 12 | 0.24 | 0.52 |

| 7 | 12 | 0.24 | 0.52 |

| 8 | 14 | 0.28 | 0.80 |

| 8 | 14 | 0.28 | 0.80 |

| 9 | 7 | 0.14 | 0.94 |

| 9 | 7 | 0.14 | 0.94 |

| 10 | 3 | 0.06 | 1.00 |

| 10 | 3 | 0.06 | 1.00 |

Table 2.22

表 2.22

Find the 28th percentile. Notice the 0.28 in the "cumulative relative frequency" column. Twenty-eight percent of 50 data values is 14 values. There are 14 values less than the 28th percentile. They include the two 4s, the five 5s, and the seven 6s. The 28th percentile is between the last six and the first seven. The 28th percentile is 6.5.

求第 28 百分位数。注意"累积相对频数"列中的 0.28。50 个数据值的 28% 是 14 个值。小于第 28 百分位数的有 14 个值,包括两个 4、五个 5 和七个 6。第 28 百分位数位于最后一个 6 与第一个 7 之间。第 28 百分位数为 6.5。

Find the median. Look again at the "cumulative relative frequency" column and find 0.52. The median is the 50th percentile or the second quartile. 50% of 50 is 25. There are 25 values less than the median. They include the two 4s, the five 5s, the seven 6s, and eleven of the 7s. The median or 50th percentile is between the 25th, or seven, and 26th, or seven, values. The median is seven.

求中位数。再看"累积相对频数"列,找到 0.52。中位数是第 50 百分位数,即第二四分位数。50 的 50% 是 25。小于中位数的有 25 个值,包括两个 4、五个 5、七个 6,以及十一个 7。中位数(第 50 百分位数)位于第 25 个值(即 7)与第 26 个值(即 7)之间。中位数为 7。

Find the third quartile. The third quartile is the same as the 75th percentile. You can "eyeball" this answer. If you look at the "cumulative relative frequency" column, you find 0.52 and 0.80. When you have all the fours, fives, sixes and sevens, you have 52% of the data. When you include all the 8s, you have 80% of the data. The 75th percentile, then, must be an eight. Another way to look at the problem is to find 75% of 50, which is 37.5, and round up to 38. The third quartile, *Q*3, is the 38th value, which is an eight. You can check this answer by counting the values. (There are 37 values below the third quartile and 12 values above.)

求第三四分位数。第三四分位数等于第 75 百分位数。你可以"目测"这个答案。如果查看"累积相对频数"列,你会找到 0.52 和 0.80。当取完所有的 4、5、6 和 7 时,你已拥有 52% 的数据;当再包含所有的 8 时,你就有了 80% 的数据。因此第 75 百分位数必然为 8。另一种思路是求 50 的 75%,即 37.5,向上取整为 38。第三四分位数 *Q*3 是第 38 个值,即 8。你可以通过计数来验证这个答案(第三四分位数以下有 37 个值,以上有 12 个值)。

Forty bus drivers were asked how many hours they spend each day running their routes (rounded to the nearest hour). Find the 65th percentile.

40 名公交车司机被问及每天跑线路花多少小时(四舍五入到最近的整数)。求第 65 百分位数。

| Amount of time spent on route (hours) | Frequency | Relative Frequency | Cumulative Relative Frequency |

| 行驶路线所花时间(小时) | 频数 | 相对频数 | 累积相对频数 |

|---------------------------------------|-----------|--------------------|-------------------------------|

|---------------------------------------|-----------|--------------------|-------------------------------|

| 2 | 12 | 0.30 | 0.30 |

| 2 | 12 | 0.30 | 0.30 |

| 3 | 14 | 0.35 | 0.65 |

| 3 | 14 | 0.35 | 0.65 |

| 4 | 10 | 0.25 | 0.90 |

| 4 | 10 | 0.25 | 0.90 |

| 5 | 4 | 0.10 | 1.00 |

| 5 | 4 | 0.10 | 1.00 |

Table 2.23

表 2.23

Problem 问题

Using Table 2.22:

使用表 2.22:

1. Find the 80th percentile.

1. 求第 80 百分位数。

2. Find the 90th percentile.

2. 求第 90 百分位数。

3. Find the first quartile. What is another name for the first quartile?

3. 求第一四分位数。第一四分位数的另一个名称是什么?

Solution 解答

Using the data from the frequency table, we have:

利用频数表中的数据,我们得到:

1. The 80th percentile is between the last eight and the first nine in the table (between the 40th and 41st values). Therefore, we need to take the mean of the 40th an 41st values. The 80th percentile $= \frac{8 + 9}{2} = 8.5$

1. 第 80 百分位数位于表中最后一个 8 与第一个 9 之间(即第 40 个与第 41 个值之间)。因此,我们需要取第 40 个和第 41 个值的均值。第 80 百分位数 $= \frac{8 + 9}{2} = 8.5$

2. The 90th percentile will be the 45th data value (location is 0.90(50) = 45) and the 45th data value is nine.

2. 第 90 百分位数将是第 45 个数据值(位置为 0.90(50) = 45),而第 45 个数据值为 9。

3. *Q*1 is also the 25th percentile. The 25th percentile location calculation: *P*25 = 0.25(50) = 12.5 ≈ 13 the 13th data value. Thus, the 25th percentile is six.

3. *Q*1 也是第 25 百分位数。第 25 百分位数的位置计算如下:*P*25 = 0.25(50) = 12.5 ≈ 13,即第 13 个数据值。因此第 25 百分位数为 6。

Refer to the Table 2.23. Find the third quartile. What is another name for the third quartile?

参考表 2.23。求第三四分位数。第三四分位数的另一个名称是什么?

Your instructor or a member of the class will ask everyone in class how many sweaters they own. Answer the following questions:

你的老师或班上某位同学会询问班里每个人拥有多少件毛衣。回答下列问题:

1. How many students were surveyed?

1. 调查了多少名学生?

2. What kind of sampling did you do?

2. 你采用了哪种抽样方式?

3. Construct two different histograms. For each, starting value = \_\_\_\_\_ ending value = \_\_\_\_.

3. 制作两个不同的直方图。对每个直方图,起始值 = \_\_\_\_\_,终止值 = \_\_\_\_。

4. Find the median, first quartile, and third quartile.

4. 求中位数、第一四分位数和第三四分位数。

5. Construct a table of the data to find the following:

5. 制作该数据的表格,以求出下列各项:

1. the 10th percentile

1. 第 10 百分位数

2. the 70th percentile

2. 第 70 百分位数

3. the percent of students who own less than four sweaters

3. 拥有少于四件毛衣的学生所占的百分比

A Formula for Finding the *k*th Percentile 求第 k 百分位数的公式

If you were to do a little research, you would find several formulas for calculating the *k*th percentile. Here is one of them.

如果你做一点研究,就会发现计算第 k 百分位数的公式有好几种。下面是其中之一。

*k* = the *kth* percentile. It may or may not be part of the data.

*k* = 第 *k*th 百分位数。它可能是、也可能不是数据的一部分。

*i* = the index (ranking or position of a data value)

*i* = 索引(某个数据值的排名或位置)

*n* = the total number of data

*n* = 数据的总个数

Problem 问题

Listed are 29 ages for Academy Award winning best actors *in order from smallest to largest.*

下面按从小到大顺序排列了 29 位奥斯卡最佳男主角获奖者的年龄。

18; 21; 22; 25; 26; 27; 29; 30; 31; 33; 36; 37; 41; 42; 47; 52; 55; 57; 58; 62; 64; 67; 69; 71; 72; 73; 74; 76; 77

18;21;22;25;26;27;29;30;31;33;36;37;41;42;47;52;55;57;58;62;64;67;69;71;72;73;74;76;77

1. Find the 70th percentile.

1. 求第 70 百分位数。

2. Find the 83rd percentile.

2. 求第 83 百分位数。

Solution 解答

1. - *k* = 70

1. - *k* = 70

*i* = $\frac{k}{100}$ (*n* + 1) = ($\frac{70}{100}$)(29 + 1) = 21. Twenty-one is an integer, and the data value in the 21st position in the ordered data set is 64. The 70th percentile is 64 years.

*i* = $\frac{k}{100}$ (*n* + 1) = ($\frac{70}{100}$)(29 + 1) = 21。21 是整数,有序数据集中第 21 个位置的数据值是 64。第 70 百分位数为 64 岁。

2. - *k* = 83rd percentile

2. - *k* = 第 83rd 百分位数

*i* = $\frac{k}{100}$ (*n* + 1) = ($\frac{83}{100}$)(29 + 1) = 24.9, which is NOT an integer. Round it down to 24 and up to 25. The age in the 24th position is 71 and the age in the 25th position is 72. Average 71 and 72. The 83rd percentile is 71.5 years.

*i* = $\frac{k}{100}$ (*n* + 1) = ($\frac{83}{100}$)(29 + 1) = 24.9,这不是整数。将其向下取整为 24、向上取整为 25。第 24 个位置的年龄是 71,第 25 个位置的年龄是 72。取 71 和 72 的平均数。第 83 百分位数为 71.5 岁。

Listed are 29 ages for Academy Award winning best actors *in order from smallest to largest.*

下面按从小到大顺序排列了 29 位奥斯卡最佳男主角获奖者的年龄。

18; 21; 22; 25; 26; 27; 29; 30; 31; 33; 36; 37; 41; 42; 47; 52; 55; 57; 58; 62; 64; 67; 69; 71; 72; 73; 74; 76; 77

18;21;22;25;26;27;29;30;31;33;36;37;41;42;47;52;55;57;58;62;64;67;69;71;72;73;74;76;77

Calculate the 20th percentile and the 55th percentile.

计算第 20 百分位数和第 55 百分位数。

You can calculate percentiles using calculators and computers. There are a variety of online calculators.

你可以使用计算器或计算机来计算百分位数。网上有多种在线计算器。

A Formula for Finding the Percentile of a Value in a Data Set 求数据集中某一数值百分位数的公式

Problem 问题

Listed are 29 ages for Academy Award winning best actors *in order from smallest to largest.*

下面按从小到大顺序排列了 29 位奥斯卡最佳男主角获奖者的年龄。

18; 21; 22; 25; 26; 27; 29; 30; 31; 33; 36; 37; 41; 42; 47; 52; 55; 57; 58; 62; 64; 67; 69; 71; 72; 73; 74; 76; 77

18;21;22;25;26;27;29;30;31;33;36;37;41;42;47;52;55;57;58;62;64;67;69;71;72;73;74;76;77

1. Find the percentile for 58.

1. 求 58 的百分位数。

2. Find the percentile for 25.

2. 求 25 的百分位数。

Solution 解答

1. Counting from the bottom of the list, there are 18 data values less than 58. There is one value of 58.

1. 从列表底部往上数,小于 58 的数据值有 18 个。58 这个值出现了 1 次。

*x* = 18 and *y* = 1.$\frac{x + 0.5y}{n}$(100) = $\frac{18 + 0.5(1)}{29}$(100) = 63.80. 58 is the 64th percentile.

*x* = 18 且 *y* = 1.$\frac{x + 0.5y}{n}$(100) = $\frac{18 + 0.5(1)}{29}$(100) = 63.80。58 是第 64 百分位数。

2. Counting from the bottom of the list, there are three data values less than 25. There is one value of 25.

2. 从列表底部往上数,小于 25 的数据值有 3 个。25 这个值出现了 1 次。

*x* = 3 and *y* = 1.$\frac{x + 0.5y}{n}$(100) = $\frac{3 + 0.5(1)}{29}$(100) = 12.07. Twenty-five is the 12th percentile.

*x* = 3 且 *y* = 1.$\frac{x + 0.5y}{n}$(100) = $\frac{3 + 0.5(1)}{29}$(100) = 12.07。25 是第 12 百分位数。

Listed are 30 ages for Academy Award winning best actors in order from smallest to largest.

下面按从小到大顺序排列了 30 位奥斯卡最佳男主角获奖者的年龄。

18; 21; 22; 25; 26; 27; 29; 30; 31, 31; 33; 36; 37; 41; 42; 47; 52; 55; 57; 58; 62; 64; 67; 69; 71; 72; 73; 74; 76; 77

18;21;22;25;26;27;29;30;31, 31;33;36;37;41;42;47;52;55;57;58;62;64;67;69;71;72;73;74;76;77

Find the percentiles for 47 and 31.

求 47 和 31 的百分位数。

Interpreting Percentiles, Quartiles, and Median 解释百分位数、四分位数与中位数

A percentile indicates the relative standing of a data value when data are sorted into numerical order from smallest to largest. Percentages of data values are less than or equal to the pth percentile. For example, 15% of data values are less than or equal to the 15th percentile.

百分位数表示当数据按数值从小到大排序时,某一数据值的相对位置。小于或等于第 p 百分位数的数据值占数据值的百分比。例如,15% 的数据值小于或等于第 15 百分位数。

A percentile may or may not correspond to a value judgment about whether it is "good" or "bad." The interpretation of whether a certain percentile is "good" or "bad" depends on the context of the situation to which the data applies. In some situations, a low percentile would be considered "good;" in other contexts a high percentile might be considered "good". In many situations, there is no value judgment that applies.

百分位数未必对应于关于它"好"或"坏"的价值判断。某个百分位数是否"好"或"坏"的解释,取决于数据所处情境的背景。在某些情境下,较低的百分位数会被认为"好";在另一些情境下,较高的百分位数可能被认为"好"。在许多情境下,并不存在适用的价值判断。

Understanding how to interpret percentiles properly is important not only when describing data, but also when calculating probabilities in later chapters of this text.

正确理解如何解释百分位数,不仅在描述数据时很重要,在本书后面章节计算概率时也很重要。

When writing the interpretation of a percentile in the context of the given data, the sentence should contain the following information.

在结合给定数据背景写百分位数的解释时,句子应包含下列信息。

Problem 问题

On a timed math test, the first quartile for time it took to finish the exam was 35 minutes. Interpret the first quartile in the context of this situation.

在一次限时的数学测验中,完成考试所需时间的第一四分位数为 35 分钟。结合这一情境解释第一四分位数。

Solution 解答

For the 100-meter dash, the third quartile for times for finishing the race was 11.5 seconds. Interpret the third quartile in the context of the situation.

在 100 米短跑中,完成比赛所用时间的第三四分位数为 11.5 秒。结合这一情境解释第三四分位数。

Problem 问题

On a 20 question math test, the 70th percentile for number of correct answers was 16. Interpret the 70th percentile in the context of this situation.

在一份 20 道题的数学测验中,答对题数的第 70 百分位数为 16。结合这一情境解释第 70 百分位数。

On a 60 point written assignment, the 80th percentile for the number of points earned was 49. Interpret the 80th percentile in the context of this situation.

在一份 60 分的书面作业中,得分数的第 80 百分位数为 49。结合这一情境解释第 80 百分位数。

Problem 问题

At a community college, it was found that the 30th percentile of credit units that students are enrolled for is seven units. Interpret the 30th percentile in the context of this situation.

在一所社区学院中,发现学生所修学分单元数的第 30 百分位数为 7 个单元。结合这一情境解释第 30 百分位数。

During a season, the 40th percentile for points scored per player in a game is eight. Interpret the 40th percentile in the context of this situation.

在某个赛季中,每名球员单场得分的第 40 百分位数为 8 分。结合这一情境解释第 40 百分位数。

Sharpe Middle School is applying for a grant that will be used to add fitness equipment to the gym. The principal surveyed 15 anonymous students to determine how many minutes a day the students spend exercising. The results from the 15 anonymous students are shown.

Sharpe 中学正在申请一笔用于为体育馆添置健身器材的补助金。校长调查了 15 名匿名学生,以确定他们每天锻炼多少分钟。这 15 名匿名学生的结果如下所示。

0 minutes; 40 minutes; 60 minutes; 30 minutes; 60 minutes

0 分钟;40 分钟;60 分钟;30 分钟;60 分钟

10 minutes; 45 minutes; 30 minutes; 300 minutes; 90 minutes;

10 分钟;45 分钟;30 分钟;300 分钟;90 分钟;

30 minutes; 120 minutes; 60 minutes; 0 minutes; 20 minutes

30 分钟;120 分钟;60 分钟;0 分钟;20 分钟

Determine the following five values.

确定下列五个数值。

If you were the principal, would you be justified in purchasing new fitness equipment? Since 75% of the students exercise for 60 minutes or less daily, and since the *IQR* is 40 minutes (60 – 20 = 40), we know that half of the students surveyed exercise between 20 minutes and 60 minutes daily. This seems a reasonable amount of time spent exercising, so the principal would be justified in purchasing the new equipment.

如果你是校长,是否有理由购置新的健身器材?由于 75% 的学生每天锻炼 60 分钟或更少,且 IQR 为 40 分钟(60 – 20 = 40),我们知道被调查的学生中有一半每天锻炼 20 到 60 分钟。这似乎是合理的锻炼时长,因此校长有理由购置新器材。

However, the principal needs to be careful. The value 300 appears to be a potential outlier.

不过,校长需要谨慎。数值 300 看起来是一个潜在的离群值。

*Q*3 + 1.5(*IQR*) = 60 + (1.5)(40) = 120.

*Q*3 + 1.5(*IQR*) = 60 + (1.5)(40) = 120.

The value 300 is greater than 120 so it is a potential outlier. If we delete it and calculate the five values, we get the following values:

数值 300 大于 120,因此它是一个潜在的离群值。如果我们将其删除并重新计算这五个数值,会得到以下结果:

We still have 75% of the students exercising for 60 minutes or less daily and half of the students exercising between 20 and 60 minutes a day. However, 15 students is a small sample and the principal should survey more students to be sure of his survey results.

我们仍然有 75% 的学生每天锻炼 60 分钟或更少,且一半学生每天锻炼 20 到 60 分钟。然而,15 名学生是一个小样本,校长应当调查更多学生,以确认其调查结果。

2.4 Box Plots 2.4 箱线图

Box plots (also called box-and-whisker plots or box-whisker plots) give a good graphical image of the concentration of the data. They also show how far the extreme values are from most of the data. A box plot is constructed from five values: the minimum value, the first quartile, the median, the third quartile, and the maximum value. We use these values to compare how close other data values are to them.

箱线图(又称箱须图或箱-须图)能很好地以图形展示数据的集中情况,也能显示极端值距离大部分数据有多远。箱线图由五个数值构成:最小值、第一四分位数、中位数、第三四分位数和最大值。我们利用这些值来比较其他数据值与之的接近程度。

To construct a box plot, use a horizontal or vertical number line and a rectangular box. The smallest and largest data values label the endpoints of the axis. The first quartile marks one end of the box and the third quartile marks the other end of the box. Approximately the middle 50 percent of the data fall inside the box. The "whiskers" extend from the ends of the box to the smallest and largest data values. The median or second quartile can be between the first and third quartiles, or it can be one, or the other, or both. The box plot gives a good, quick picture of the data.

绘制箱线图时,使用一条水平或竖直的数轴和一个矩形箱。最小和最大的数据值标在数轴的两端。第一四分位数标记箱的一端,第三四分位数标记另一端。大约中间 50% 的数据落在箱内。"须"从箱的两端延伸到最小和最大的数据值。中位数(第二四分位数)可以在第一与第三四分位数之间,也可以在其中某一端,或同时位于两端。箱线图能快速、清晰地呈现数据。

You may encounter box-and-whisker plots that have dots marking outlier values. In those cases, the whiskers are not extending to the minimum and maximum values.

你可能会遇到用圆点标出离群值的箱须图。此时,须不会延伸到最小值和最大值。

Consider, again, this dataset.

再次考虑这个数据集。

1; 1; 2; 2; 4; 6; 6.8; 7.2; 8; 8.3; 9; 10; 10; 11.5

1;1;2;2;4;6;6.8;7.2;8;8.3;9;10;10;11.5

The first quartile is two, the median is seven, and the third quartile is nine. The smallest value is one, and the largest value is 11.5. The following image shows the constructed box plot.

第一四分位数是 2,中位数是 7,第三四分位数是 9。最小值是 1,最大值是 11.5。下图展示了绘制好的箱线图。

See the calculator instructions on the TI web site or in the appendix.

参见 TI 网站或附录中的计算器操作说明。

The two whiskers extend from the first quartile to the smallest value and from the third quartile to the largest value. The median is shown with a dashed line.

两条须分别从第一四分位数延伸到最小值,以及从第三四分位数延伸到最大值。中位数用虚线表示。

It is important to start a box plot with a scaled number line. Otherwise the box plot may not be useful.

绘制箱线图时必须先使用一条带刻度的数轴,否则箱线图可能失去参考价值。

The following data are the heights of 40 students in a statistics class.

以下数据是某统计课上 40 名学生的身高。

59; 60; 61; 62; 62; 63; 63; 64; 64; 64; 65; 65; 65; 65; 65; 65; 65; 65; 65; 66; 66; 67; 67; 68; 68; 69; 70; 70; 70; 70; 70; 71; 71; 72; 72; 73; 74; 74; 75; 77

59;60;61;62;62;63;63;64;64;64;65;65;65;65;65;65;65;65;65;66;66;67;67;68;68;69;70;70;70;70;70;71;71;72;72;73;74;74;75;77

Construct a box plot with the following properties; the calculator intructions for the minimum and maximum values as well as the quartiles follow the example.

根据以下特征绘制箱线图;关于最小值、最大值以及各四分位数的计算器操作步骤见示例。

1. Each quarter has approximately 25% of the data.

1. 每个四分位大约包含 25% 的数据。

2. The spreads of the four quarters are 64.5 – 59 = 5.5 (first quarter), 66 – 64.5 = 1.5 (second quarter), 70 – 66 = 4 (third quarter), and 77 – 70 = 7 (fourth quarter). So, the second quarter has the smallest spread and the fourth quarter has the largest spread.

2. 四个四分位的跨度分别为 64.5 – 59 = 5.5(第一四分位)、66 – 64.5 = 1.5(第二四分位)、70 – 66 = 4(第三四分位)、77 – 70 = 7(第四四分位)。因此,第二四分位跨度最小,第四四分位跨度最大。

3. Range = maximum value – the minimum value = 77 – 59 = 18

3. 极差 = 最大值 – 最小值 = 77 – 59 = 18

4. Interquartile Range: *IQR* = *Q*3 – *Q*1 = 70 – 64.5 = 5.5.

4. 四分位距:*IQR* = *Q*3 – *Q*1 = 70 – 64.5 = 5.5。

5. The interval 59–65 has more than 25% of the data so it has more data in it than the interval 66 through 70 which has 25% of the data.

5. 区间 59–65 包含超过 25% 的数据,因此其中数据比含有 25% 数据的区间 66 至 70 更多。

6. The middle 50% (middle half) of the data has a range of 5.5 inches.

6. 数据的中间 50%(中间一半)的极差为 5.5 英寸。

To find the minimum, maximum, and quartiles:

求最小值、最大值和四分位数:

Enter data into the list editor (Pres STAT 1:EDIT). If you need to clear the list, arrow up to the name L1, press CLEAR, and then arrow down.

将数据输入列表编辑器(按 STAT 1:EDIT)。如需清除列表,向上移到 L1 名称,按 CLEAR,再向下移。

Put the data values into the list L1.

将数据值输入列表 L1。

Press STAT and arrow to CALC. Press 1:1-VarStats. Enter L1.

按 STAT,移到 CALC,按 1:1-VarStats,输入 L1。

Press ENTER.

按 ENTER。

Use the down and up arrow keys to scroll.

使用向下和向上方向键滚动查看。

Smallest value = 59.

最小值 = 59。

Largest value = 77.

最大值 = 77。

*Q*1: First quartile = 64.5.

*Q*1:第一四分位数 = 64.5。

*Q*2: Second quartile or median = 66.

*Q*2:第二四分位数或中位数 = 66。

*Q*3: Third quartile = 70.

*Q*3:第三四分位数 = 70。

To construct the box plot:

绘制箱线图:

Press 4:Plotsoff. Press ENTER.

按 4:Plotsoff,按 ENTER。

Arrow down and then use the right arrow key to go to the fifth picture, which is the box plot. Press ENTER.

向下移,再用右方向键移到第五个图形(即箱线图),按 ENTER。

Arrow down to Xlist: Press 2nd 1 for L1

向下移到 Xlist:按 2nd 1 选择 L1。

Arrow down to Freq: Press ALPHA. Press 1.

向下移到 Freq:按 ALPHA,再按 1。

Press Zoom. Press 9: ZoomStat.

按 Zoom,按 9: ZoomStat。

Press TRACE, and use the arrow keys to examine the box plot.

按 TRACE,用方向键查看箱线图。

The following data are the number of pages in 40 books on a shelf. Construct a box plot using a graphing calculator, and state the interquartile range.

以下是书架上 40 本书的页数。用图形计算器绘制箱线图,并说明四分位距。

136; 140; 178; 190; 205; 215; 217; 218; 232; 234; 240; 255; 270; 275; 290; 301; 303; 315; 317; 318; 326; 333; 343; 349; 360; 369; 377; 388; 391; 392; 398; 400; 402; 405; 408; 422; 429; 450; 475; 512

136;140;178;190;205;215;217;218;232;234;240;255;270;275;290;301;303;315;317;318;326;333;343;349;360;369;377;388;391;392;398;400;402;405;408;422;429;450;475;512

For some sets of data, some of the largest value, smallest value, first quartile, median, and third quartile may be the same. For instance, you might have a data set in which the median and the third quartile are the same. In this case, the diagram would not have a dotted line inside the box displaying the median. The right side of the box would display both the third quartile and the median. For example, if the smallest value and the first quartile were both one, the median and the third quartile were both five, and the largest value was seven, the box plot would look like:

对某些数据集,最大值、最小值、第一四分位数、中位数和第三四分位数中可能有部分相同。例如,某个数据集的中位数与第三四分位数相同。此时图中箱内不会出现表示中位数的虚线,箱的右侧会同时显示第三四分位数和中位数。举例来说,若最小值与第一四分位数都是 1,中位数与第三四分位数都是 5,最大值为 7,则箱线图会如下所示:

In this case, at least 25% of the values are equal to one. Twenty-five percent of the values are between one and five, inclusive. At least 25% of the values are equal to five. The top 25% of the values fall between five and seven, inclusive.

此时,至少 25% 的值等于 1。有 25% 的值介于 1 与 5 之间(含端点)。至少 25% 的值等于 5。最高的 25% 的值落在 5 与 7 之间(含端点)。

Test scores for a college statistics class held during the day are:

某大学统计课日间班的考试成绩为:

99; 56; 78; 55.5; 32; 90; 80; 81; 56; 59; 45; 77; 84.5; 84; 70; 72; 68; 32; 79; 90

99;56;78;55.5;32;90;80;81;56;59;45;77;84.5;84;70;72;68;32;79;90

Test scores for a college statistics class held during the evening are:

某大学统计课夜间班的考试成绩为:

98; 78; 68; 83; 81; 89; 88; 76; 65; 45; 98; 90; 80; 84.5; 85; 79; 78; 98; 90; 79; 81; 25.5

98;78;68;83;81;89;88;76;65;45;98;90;80;84.5;85;79;78;98;90;79;81;25.5

Problem 问题

1. Find the smallest and largest values, the median, and the first and third quartile for the day class.

1. 求日间班的最小值、最大值、中位数以及第一、第三四分位数。

2. Find the smallest and largest values, the median, and the first and third quartile for the night class.

2. 求夜间班的最小值、最大值、中位数以及第一、第三四分位数。

3. For each data set, what percentage of the data is between the smallest value and the first quartile? the first quartile and the median? the median and the third quartile? the third quartile and the largest value? What percentage of the data is between the first quartile and the largest value?

3. 对每个数据集,有多少百分比的数据位于最小值与第一四分位数之间?第一四分位数与中位数之间?中位数与第三四分位数之间?第三四分位数与最大值之间?又有多少百分比的数据位于第一四分位数与最大值之间?

4. Create a box plot for each set of data. Use one number line for both box plots.

4. 为每组数据各绘制一个箱线图,两条箱线图使用同一条数轴。

5. Which box plot has the widest spread for the middle 50% of the data (the data between the first and third quartiles)? What does this mean for that set of data in comparison to the other set of data?

5. 哪个箱线图的中间 50% 数据(第一与第三四分位数之间的数据)跨度最大?相对于另一组数据,这意味着什么?

Solution 解答

1. - Min = 32

1. - 最小值 = 32

2. - Min = 25.5

2. - 最小值 = 25.5

3. Day class: There are six data values ranging from 32 to 56: 30%. There are six data values ranging from 56 to 74.5: 30%. There are five data values ranging from 74.5 to 82.5: 25%. There are five data values ranging from 82.5 to 99: 25%. There are 16 data values between the first quartile, 56, and the largest value, 99: 75%. Night class:

3. 日间班:在 32 到 56 之间有 6 个数据值:30%;在 56 到 74.5 之间有 6 个:30%;在 74.5 到 82.5 之间有 5 个:25%;在 82.5 到 99 之间有 5 个:25%;在第一四分位数 56 与最大值 99 之间有 16 个数据值:75%。夜间班:

4.

4. (此处应绘制箱线图)

5. The first data set has the wider spread for the middle 50% of the data. The *IQR* for the first data set is greater than the *IQR* for the second set. This means that there is more variability in the middle 50% of the first data set.

5. 第一组数据的中间 50% 数据跨度更大。第一组的 *IQR* 大于第二组。这意味着第一组中间 50% 的数据变异性更大。

The following data set shows the heights in inches for the boys in a class of 40 students.

以下数据集是某 40 人班级中男生的身高(英寸)。

66; 66; 67; 67; 68; 68; 68; 68; 68; 69; 69; 69; 70; 71; 72; 72; 72; 73; 73; 74

66;66;67;67;68;68;68;68;68;69;69;69;70;71;72;72;72;73;73;74

The following data set shows the heights in inches for the girls in a class of 40 students.

以下数据集是某 40 人班级中女生的身高(英寸)。

61; 61; 62; 62; 63; 63; 63; 65; 65; 65; 66; 66; 66; 67; 68; 68; 68; 69; 69; 69

61;61;62;62;63;63;63;65;65;65;66;66;66;67;68;68;68;69;69;69

Construct a box plot using a graphing calculator for each data set, and state which box plot has the wider spread for the middle 50% of the data.

用图形计算器为每组数据各绘制一个箱线图,并说明哪个箱线图的中间 50% 数据跨度更大。

Graph a box-and-whisker plot for the data values shown.

为所示数据值绘制一个箱须图。

10; 10; 10; 15; 35; 75; 90; 95; 100; 175; 420; 490; 515; 515; 790

10;10;10;15;35;75;90;95;100;175;420;490;515;515;790

The five numbers used to create a box-and-whisker plot are:

用于绘制箱须图的五个数值为:

The following graph shows the box-and-whisker plot.

下图展示了该箱须图。

Follow the steps you used to graph a box-and-whisker plot for the data values shown.

仿照上述步骤,为所示数据值绘制箱须图。

0; 5; 5; 15; 30; 30; 45; 50; 50; 60; 75; 110; 140; 240; 330

0;5;5;15;30;30;45;50;50;60;75;110;140;240;330

---

(分隔线)

2.5 Measures of the Center of the Data 2.5 数据中心位置的度量

The "center" of a data set is also a way of describing location. The two most widely used measures of the "center" of the data are the mean (average) and the median. To calculate the mean weight of 50 people, add the 50 weights together and divide by 50. To find the median weight of the 50 people, order the data and find the number that splits the data into two equal parts. The median is generally a better measure of the center when there are extreme values or outliers because it is not affected by the precise numerical values of the outliers. The mean is the most common measure of the center.

数据集的"中心"也是描述其位置的一种方式。使用最广泛的两种数据"中心"度量是均值(平均数)和中位数。要计算 50 人的平均体重,把 50 个体重相加再除以 50。要找这 50 人的体重中位数,需将数据排序,找出将数据分为两等份的数。当存在极端值或离群值时,中位数通常是更好的中心度量,因为它不受离群值具体数值的影响。均值是使用最广泛的中心度量。

The words “mean” and “average” are often used interchangeably. The substitution of one word for the other is common practice. The technical term is “arithmetic mean” and “average” is technically a center location. However, in practice among non-statisticians, “average" is commonly accepted for “arithmetic mean.”

"mean"(均值)和"average"(平均数)经常互换使用,一词替代另一词是常见做法。技术术语是"算术平均"(arithmetic mean),而"average"严格来说是一种中心位置。然而,在非统计学家群体的实践中,"average"通常被接受为"算术平均"的同义用法。

When each value in the data set is not unique, the mean can be calculated by multiplying each distinct value by its frequency and then dividing the sum by the total number of data values. The letter used to represent the sample mean is an *x* with a bar over it (pronounced “*x* bar”): $\overset{–}{x}$.

当数据集中的各个取值并非唯一时,可将每个不同取值乘以其频数,再将总和除以数据值总个数来计算均值。表示样本均值的字母是上方带横线的 *x*(读作"x bar"):$\overset{–}{x}$。

The Greek letter *μ* (pronounced "mew") represents the population mean. One of the requirements for the sample mean to be a good estimate of the population mean is for the sample taken to be truly random.

希腊字母 *μ*(读作"mew")表示总体均值。要使样本均值成为总体均值的一个良好估计,所抽取的样本必须真正随机,这是其要求之一。

To see that both ways of calculating the mean are the same, consider the sample:

为说明两种计算均值的方式结果相同,考虑如下样本:

1; 1; 1; 2; 2; 3; 4; 4; 4; 4; 4

1;1;1;2;2;3;4;4;4;4;4

$$\overline{x} = \frac{1 + 1 + 1 + 2 + 2 + 3 + 4 + 4 + 4 + 4 + 4}{11} = 2.7$$ $$\overline{x} = \frac{3(1) + 2(2) + 1(3) + 5(4)}{11} = 2.7$$

$$\overline{x} = \frac{1 + 1 + 1 + 2 + 2 + 3 + 4 + 4 + 4 + 4 + 4}{11} = 2.7$$ $$\overline{x} = \frac{3(1) + 2(2) + 1(3) + 5(4)}{11} = 2.7$$

In the second calculation, the frequencies are 3, 2, 1, and 5.

在第二种计算中,频数分别为 3、2、1 和 5。

You can quickly find the location of the median by using the expression $\frac{n + 1}{2}$.

利用表达式 $\frac{n + 1}{2}$ 可快速确定中位数的位置。

The letter *n* is the total number of data values in the sample. If *n* is an odd number, the median is the middle value of the ordered data (ordered smallest to largest). If *n* is an even number, the median is equal to the two middle values added together and divided by two after the data has been ordered. For example, if the total number of data values is 97, then $\frac{n + 1}{2}$= $\frac{97 + 1}{2}$ = 49. The median is the 49th value in the ordered data. If the total number of data values is 100, then $\frac{n + 1}{2}$= $\frac{100 + 1}{2}$ = 50.5. The median occurs midway between the 50th and 51st values. The location of the median and the value of the median are not the same. The upper case letter *M* is often used to represent the median. The next example illustrates the location of the median and the value of the median.

字母 *n* 是样本中数据值的总个数。若 *n* 为奇数,中位数就是有序数据(从小到大排序)的中间值。若 *n* 为偶数,中位数等于数据排序后两个中间值相加再除以 2。例如,若数据值总数为 97,则 $\frac{n + 1}{2}$= $\frac{97 + 1}{2}$ = 49,中位数是有序数据中的第 49 个值。若数据值总数为 100,则 $\frac{n + 1}{2}$= $\frac{100 + 1}{2}$ = 50.5,中位数出现在第 50 个与第 51 个值的正中间。中位数的位置与中位数的数值并不相同。大写字母 *M* 常用来表示中位数。下一个例子说明了中位数位置与中位数数值的区别。

Problem 问题

AIDS data indicating the number of months a patient with AIDS lives after taking a new antibody drug are as follows (smallest to largest):

下列 AIDS 数据显示患者在服用一种新抗体药物后存活的月数(从小到大):

3; 4; 8; 8; 10; 11; 12; 13; 14; 15; 15; 16; 16; 17; 17; 18; 21; 22; 22; 24; 24; 25; 26; 26; 27; 27; 29; 29; 31; 32; 33; 33; 34; 34; 35; 37; 40; 44; 44; 47;

3;4;8;8;10;11;12;13;14;15;15;16;16;17;17;18;21;22;22;24;24;25;26;26;27;27;29;29;31;32;33;33;34;34;35;37;40;44;44;47

Calculate the mean and the median.

计算均值与中位数。

Solution 解答

The calculation for the mean is:

均值的计算如下:

$\overline{x} = \frac{\left\lbrack 3 + 4 + (8)(2) + 10 + 11 + 12 + 13 + 14 + (15)(2) + (16)(2) + \text{...} + 35 + 37 + 40 + (44)(2) + 47 \right\rbrack}{40} = {23.6}$

$\overline{x} = \frac{\left\lbrack 3 + 4 + (8)(2) + 10 + 11 + 12 + 13 + 14 + (15)(2) + (16)(2) + \text{...} + 35 + 37 + 40 + (44)(2) + 47 \right\rbrack}{40} = {23.6}$

To find the median, *M*, first use the formula for the location. The location is:

求中位数 *M*,先使用位置公式。位置为:

$\frac{n + 1}{2} = \frac{40 + 1}{2} = 20.5$

$\frac{n + 1}{2} = \frac{40 + 1}{2} = 20.5$

Starting at the smallest value, the median is located between the 20th and 21st values (the two 24s):

从最小值开始数,中位数位于第 20 个与第 21 个值(两个 24)之间:

3; 4; 8; 8; 10; 11; 12; 13; 14; 15; 15; 16; 16; 17; 17; 18; 21; 22; 22; 24; 24; 25; 26; 26; 27; 27; 29; 29; 31; 32; 33; 33; 34; 34; 35; 37; 40; 44; 44; 47;

3;4;8;8;10;11;12;13;14;15;15;16;16;17;17;18;21;22;22;24;24;25;26;26;27;27;29;29;31;32;33;33;34;34;35;37;40;44;44;47

$M = \frac{24 + 24}{2} = 24$

$M = \frac{24 + 24}{2} = 24$

To find the mean and the median:

求均值与中位数:

Clear list L1. Pres STAT 4:ClrList. Enter 2nd 1 for list L1. Press ENTER.

清除列表 L1。按 STAT 4:ClrList,输入 2nd 1 选择列表 L1,按 ENTER。

Enter data into the list editor. Press STAT 1:EDIT.

将数据输入列表编辑器。按 STAT 1:EDIT。

Put the data values into list L1.

将数据值输入列表 L1。

Press STAT and arrow to CALC. Press 1:1-VarStats. Press 2nd 1 for L1 and then ENTER.

按 STAT,移到 CALC,按 1:1-VarStats,输入 2nd 1 选择 L1,再按 ENTER。

Press the down and up arrow keys to scroll.

按向下和向上方向键滚动查看。

$\overline{x}$ = 23.6, *M* = 24

$\overline{x}$ = 23.6,*M* = 24

The following data show the number of months patients typically wait on a transplant list before getting surgery. The data are ordered from smallest to largest. Calculate the mean and median.

以下数据显示患者在进入移植名单后、手术前通常等待的月数。数据已从小到大排序。计算均值与中位数。

3; 4; 5; 7; 7; 7; 7; 8; 8; 9; 9; 10; 10; 10; 10; 10; 11; 12; 12; 13; 14; 14; 15; 15; 17; 17; 18; 19; 19; 19; 21; 21; 22; 22; 23; 24; 24; 24; 24

3;4;5;7;7;7;7;8;8;9;9;10;10;10;10;10;11;12;12;13;14;14;15;15;17;17;18;19;19;19;21;21;22;22;23;24;24;24;24

Problem 问题

Suppose that in a small town of 50 people, one person earns \$5,000,000 per year and the other 49 each earn \$30,000. Which is the better measure of the "center": the mean or the median?

假设在一个 50 人的小镇上,一人年薪 \$5,000,000,其余 49 人各年薪 \$30,000。"中心"的更好度量是均值还是中位数?

Solution 解答

$\overline{x} = \frac{5,000,000 + 49(30,000)}{50} = 129,400$

$\overline{x} = \frac{5,000,000 + 49(30,000)}{50} = 129,400$

*M* = 30,000

*M* = 30,000

(There are 49 people who earn \$30,000 and one person who earns \$5,000,000.)

(有 49 人年薪 \$30,000,一人年薪 \$5,000,000。)

The median is a better measure of the "center" than the mean because 49 of the values are 30,000 and one is 5,000,000. The 5,000,000 is an outlier. The 30,000 gives us a better sense of the middle of the data.

中位数比均值是更好的"中心"度量,因为 49 个值是 30,000,而一个是 5,000,000。5,000,000 是一个离群值,30,000 能更好地反映数据的中间位置。

In a sample of 60 households, one house is worth \$2,500,000. Twenty-nine houses are worth \$280,000, and all the others are worth \$315,000. Which is the better measure of the "center": the mean or the median?

在一个 60 户家庭的样本中,有一户价值 \$2,500,000,29 户价值 \$280,000,其余各户价值 \$315,000。"中心"的更好度量是均值还是中位数?

Another measure of the center is the mode. The mode is the most frequent value. There can be more than one mode in a data set as long as those values have the same frequency and that frequency is the highest. A data set with two modes is called bimodal.

中心的另一种度量是众数。众数就是出现次数最多的值。只要某些值出现次数相同且为最高频数,一个数据集就可以有多个众数。有两个众数的数据集称为双众数(bimodal)。

Statistics exam scores for 20 students are as follows:

20 名学生的统计考试成绩如下:

50; 53; 59; 59; 63; 63; 72; 72; 72; 72; 72; 76; 78; 81; 83; 84; 84; 84; 90; 93

50;53;59;59;63;63;72;72;72;72;72;76;78;81;83;84;84;84;90;93

Problem 问题

Find the mode.

求众数。

Solution 解答

The most frequent score is 72, which occurs five times. Mode = 72.

出现次数最多的分数是 72,共出现 5 次。众数 = 72。

The number of books checked out from the library from 25 students are as follows:

25 名学生从图书馆借出的图书册数如下:

0; 0; 0; 1; 2; 3; 3; 4; 4; 5; 5; 7; 7; 7; 7; 8; 8; 8; 9; 10; 10; 11; 11; 12; 12

0;0;0;1;2;3;3;4;4;5;5;7;7;7;7;8;8;8;9;10;10;11;11;12;12

Find the mode.

求众数。

Five real estate exam scores are 430, 430, 480, 480, 495. The data set is bimodal because the scores 430 and 480 each occur twice.

五个房地产考试成绩为 430、430、480、480、495。该数据集为双众数,因为分数 430 和 480 各出现两次。

When is the mode the best measure of the "center"? Consider a weight loss program that advertises a mean weight loss of six pounds the first week of the program. The mode might indicate that most people lose two pounds the first week, making the program less appealing.

何时众数才是"中心"的最佳度量?设想一个减肥项目,其宣传称第一周平均减重 6 磅。而众数可能表明大多数人在第一周只减了 2 磅,从而使该项目不那么吸引人。

The mode can be calculated for qualitative data as well as for quantitative data. For example, if the data set is: red, red, red, green, green, yellow, purple, black, blue, the mode is red.

众数既可用于定量数据,也可用于定性数据。例如,若数据集为:red、red、red、green、green、yellow、purple、black、blue,则众数为 red(红)。

Statistical software will easily calculate the mean, the median, and the mode. Some graphing calculators can also make these calculations. In the real world, people make these calculations using software.

统计软件可以轻松计算均值、中位数和众数。某些图形计算器也能进行这些计算。在现实中,人们借助软件来完成这些计算。

Five credit scores are 680, 680, 700, 720, 720. The data set is bimodal because the scores 680 and 720 each occur twice. Consider the annual earnings of workers at a factory. The mode is \$25,000 and occurs 150 times out of 301. The median is \$50,000 and the mean is \$47,500. What would be the best measure of the "center"?

五个信用评分为 680、680、700、720、720。该数据集为双众数,因为分数 680 和 720 各出现两次。考虑某工厂工人的年收入:众数为 \$25,000,在 301 人中出现了 150 次;中位数为 \$50,000,均值为 \$47,500。什么是"中心"的最佳度量?

The Law of Large Numbers and the Mean 大数定律与均值

The Law of Large Numbers says that if you take samples of larger and larger size from any population, then the mean $\overline{x}$ of the sample is very likely to get closer and closer to *µ*. This is discussed in more detail later in the text.

大数定律指出,若从任意总体中抽取越来越大的样本,则样本均值 $\overline{x}$ 极有可能越来越接近 *µ*。这一点将在后文更详细地讨论。

Sampling Distributions and Statistic of a Sampling Distribution 抽样分布与抽样分布的统计量

You can think of a sampling distribution as a relative frequency distribution with a great many samples. (See Sampling and Data for a review of relative frequency). Suppose thirty randomly selected students were asked the number of movies they watched the previous week. The results are in the relative frequency table shown below.

可以把抽样分布看作具有大量样本的相对频数分布。(关于相对频数,参见抽样与数据复习。)假设随机选取 30 名学生,询问他们上周观看的电影数量。结果如下表所示的相对频数表。
\# of moviesRelative Frequency
0$\frac{5}{30}$
1$\frac{15}{30}$
2$\frac{6}{30}$
3$\frac{3}{30}$
4$\frac{1}{30}$
电影数量相对频数
0$\frac{5}{30}$
1$\frac{15}{30}$
2$\frac{6}{30}$
3$\frac{3}{30}$
4$\frac{1}{30}$

Table 2.24

表 2.24

If you let the number of samples get very large (say, 300 million or more), the relative frequency table becomes a relative frequency distribution.

若让样本数量变得非常大(例如 3 亿或更多),相对频数表就会变成一个相对频数分布。

A statistic is a number calculated from a sample. Statistic examples include the mean, the median and the mode as well as others. The sample mean $\overline{x}$ is an example of a statistic which estimates the population mean *μ*.

统计量是从样本计算得到的数值。统计量的例子包括均值、中位数和众数等。样本均值 $\overline{x}$ 就是估计总体均值 *μ* 的一个统计量。

Calculating the Mean of Grouped Frequency Tables 计算分组频数表的均值

When only grouped data is available, you do not know the individual data values (we only know intervals and interval frequencies); therefore, you cannot compute an exact mean for the data set. What we must do is estimate the actual mean by calculating the mean of a frequency table. A frequency table is a data representation in which grouped data is displayed along with the corresponding frequencies. To calculate the mean from a grouped frequency table we can apply the basic definition of mean: *mean* = $\frac{data\ sum}{number\ of\ data\ values}$ We simply need to modify the definition to fit within the restrictions of a frequency table.

当只有分组数据可用时,你并不知道各个数据值(我们只知道区间和区间频数);因此,无法计算出该数据集的精确均值。我们必须做的,是通过计算频数表的均值来估计实际均值。频数表是一种数据表示方式,其中分组数据与对应的频数一起显示。要从分组频数表计算均值,我们可以套用均值的基本定义:均值 =(数据之和)/(数据值个数)。我们只需对这个定义稍作修改,以适应频数表的限制。

Since we do not know the individual data values we can instead find the midpoint of each interval. The midpoint is $\frac{lower\ boundary + upper\ boundary}{2}$. We can now modify the mean definition to be $Mean\ of\ Frequency\ Table = \frac{\sum{fm}}{\sum f}$ where *f* = the frequency of the interval and *m* = the midpoint of the interval.

由于我们不知道各个数据值,可以转而求出每个区间的中点。中点为(下边界 + 上边界)/ 2。我们现在可以把均值定义修改为:频数表均值 = Σfm / Σf,其中 f = 区间的频数,m = 区间的中点。

Problem 问题

A frequency table displaying professor Blount’s last statistic test is shown. Find the best estimate of the class mean.

下面给出一张显示 Blount 教授最后一次统计测试成绩的频数表。求班级均值的最佳估计值。

| Grade Interval | Number of Students |

| 成绩区间 | 学生人数 |

|----------------|--------------------|

|----------------|--------------------|

| 50–56.5 | 1 |

| 50–56.5 | 1 |

| 56.5–62.5 | 0 |

| 56.5–62.5 | 0 |

| 62.5–68.5 | 4 |

| 62.5–68.5 | 4 |

| 68.5–74.5 | 4 |

| 68.5–74.5 | 4 |

| 74.5–80.5 | 2 |

| 74.5–80.5 | 2 |

| 80.5–86.5 | 3 |

| 80.5–86.5 | 3 |

| 86.5–92.5 | 4 |

| 86.5–92.5 | 4 |

| 92.5–98.5 | 1 |

| 92.5–98.5 | 1 |

Table 2.25

表 2.25

Solution 解答

| Grade Interval | Midpoint |

| 成绩区间 | 中点 |

|----------------|----------|

|----------------|----------|

| 50–56.5 | 53.25 |

| 50–56.5 | 53.25 |

| 56.5–62.5 | 59.5 |

| 56.5–62.5 | 59.5 |

| 62.5–68.5 | 65.5 |

| 62.5–68.5 | 65.5 |

| 68.5–74.5 | 71.5 |

| 68.5–74.5 | 71.5 |

| 74.5–80.5 | 77.5 |

| 74.5–80.5 | 77.5 |

| 80.5–86.5 | 83.5 |

| 80.5–86.5 | 83.5 |

| 86.5–92.5 | 89.5 |

| 86.5–92.5 | 89.5 |

| 92.5–98.5 | 95.5 |

| 92.5–98.5 | 95.5 |

Table 2.26

表 2.26

$53.25(1) + 59.5(0) + 65.5(4) + 71.5(4) + 77.5(2) + 83.5(3) + 89.5(4) + 95.5(1) = 1460.25$

$53.25(1) + 59.5(0) + 65.5(4) + 71.5(4) + 77.5(2) + 83.5(3) + 89.5(4) + 95.5(1) = 1460.25$

Maris conducted a study on the effect that playing video games has on memory recall. As part of her study, she compiled the following data:

Maris 做了一项关于玩电子游戏对记忆回忆影响的研究。作为研究的一部分,她编制了以下数据:

| Hours Teenagers Spend on Video Games | Number of Teenagers |

| 青少年玩电子游戏的小时数 | 青少年人数 |

|--------------------------------------|---------------------|

|--------------------------------------|---------------------|

| 0–3.5 | 3 |

| 0–3.5 | 3 |

| 3.5–7.5 | 7 |

| 3.5–7.5 | 7 |

| 7.5–11.5 | 12 |

| 7.5–11.5 | 12 |

| 11.5–15.5 | 7 |

| 11.5–15.5 | 7 |

| 15.5–19.5 | 9 |

| 15.5–19.5 | 9 |

Table 2.27

表 2.27

What is the best estimate for the mean number of hours spent playing video games?

玩电子游戏所花小时数的均值的最佳估计值是多少?

2.6 Skewness and the Mean, Median, and Mode 2.6 偏度与均值、中位数和众数

Consider the following data set.

考虑以下数据集。

4; 5; 6; 6; 6; 7; 7; 7; 7; 7; 7; 8; 8; 8; 9; 10

4;5;6;6;6;7;7;7;7;7;7;8;8;8;9;10

This data set can be represented by following histogram. Each interval has width one, and each value is located in the middle of an interval.

该数据集可以用下面的直方图表示。每个区间的宽度为 1,每个数值位于一个区间的中间。

The histogram displays a symmetrical distribution of data. A distribution is symmetrical if a vertical line can be drawn at some point in the histogram such that the shape to the left and the right of the vertical line are mirror images of each other. The mean, the median, and the mode are each seven for these data. In a perfectly symmetrical distribution, the mean and the median are the same. This example has one mode (unimodal), and the mode is the same as the mean and median. In a symmetrical distribution that has two modes (bimodal), the two modes would be different from the mean and median.

该直方图显示数据呈对称分布。如果能在直方图的某处画一条竖直线,使得该线左右两侧的形状互为镜像,则该分布是对称的。对于这些数,均值、中位数和众数都是 7。在完全对称的分布中,均值与中位数相同。本例有一个众数(单峰),且众数等于均值和中位数。在有两个众数(双峰)的对称分布中,这两个众数会不同于均值和中位数。

The histogram for the data: 4; 5; 6; 6; 6; 7; 7; 7; 7; 8 (shown in Figure 2.17) is not symmetrical. The right-hand side seems "chopped off" compared to the left side. A distribution of this type is called skewed to the left because it is pulled out to the left.

数据 4; 5; 6; 6; 6; 7; 7; 7; 7; 8 的直方图(见图 2.17)不是对称的。与左侧相比,右侧似乎被"截断"了。这种类型的分布称为左偏,因为它被向左拖拽。

The mean is 6.3, the median is 6.5, and the mode is seven. Notice that the mean is less than the median, and they are both less than the mode. The mean and the median both reflect the skewing, but the mean reflects it more so.

均值为 6.3,中位数为 6.5,众数为 7。注意均值小于中位数,而二者都小于众数。均值和中位数都反映了偏斜,但均值反映得更明显。

The histogram for the data: 6; 7; 7; 7; 7; 8; 8; 8; 9; 10 Figure 2.18, is also not symmetrical. It is skewed to the right.

数据 6; 7; 7; 7; 7; 8; 8; 8; 9; 10 的直方图(图 2.18)也不是对称的。它是右偏的。

The mean is 7.7, the median is 7.5, and the mode is seven. Of the three statistics, the mean is the largest, while the mode is the smallest. Again, the mean reflects the skewing the most.

均值为 7.7,中位数为 7.5,众数为 7。在这三个统计量中,均值最大,而众数最小。同样,均值对偏斜的反映最明显。

The mean is affected by outliers that do not influence the mean. Therefore, when the distribution of data is skewed to the left, the mean is often less than the median. When the distribution is skewed to the right, the mean is often greater than the median. In symmetric distributions, we expect the mean and median to be approximately equal in value. This is an important connection between the shape of the distribution and the relationship of the mean and median. It is not, however, true for every data set. The most common exceptions occur in sets of discrete data.

均值会受到那些不影响均值的离群值的影响。因此,当数据分布左偏时,均值往往小于中位数;当分布右偏时,均值往往大于中位数。在对称分布中,我们预期均值与中位数近似相等。这是分布的形态与均值、中位数之间关系的一个重要联系。然而,这并非对每个数据集都成立。最常见的例外出现在离散数据集中。

Skewness and symmetry become important when we discuss probability distributions in later chapters.

当我们后面章节讨论概率分布时,偏度和对称性就变得重要了。

Problem 问题

Statistics are used to compare and sometimes identify authors. The following lists shows a simple random sample that compares the letter counts for three authors.

统计学被用来比较、有时也用来识别作者。下面列出的内容展示了一个简单随机样本,比较了三位作者的字母数。

Terry: 7; 9; 3; 3; 3; 4; 1; 3; 2; 2

Terry:7;9;3;3;3;4;1;3;2;2

Davis: 3; 3; 3; 4; 1; 4; 3; 2; 3; 1

Davis:3;3;3;4;1;4;3;2;3;1

Maris: 2; 3; 4; 4; 4; 6; 6; 6; 8; 3

Maris:2;3;4;4;4;6;6;6;8;3

1. Make a dot plot for the three authors and compare the shapes.

1. 为这三位作者制作点图,并比较其形状。

2. Calculate the mean for each.

2. 分别计算各自的均值。

3. Calculate the median for each.

3. 分别计算各自的中位数。

4. Describe any pattern you notice between the shape and the measures of center.

4. 描述你在形状与中心度量之间注意到的任何规律。

Solution 解答

1.

1.

2. Terry’s mean is 3.7, Davis’ mean is 2.7, Maris’ mean is 4.6.

2. Terry 的均值为 3.7,Davis 的均值为 2.7,Maris 的均值为 4.6。

3. Terry’s median is three, Davis’ median is three. Maris’ median is four.

3. Terry 的中位数为 3,Davis 的中位数为 3,Maris 的中位数为 4。

4. It appears that the median is always closest to the high point (the mode), while the mean tends to be farther out on the tail. In a symmetrical distribution, the mean and the median are both centrally located close to the high point of the distribution.

4. 中位数似乎总是最接近最高点(众数),而均值往往更偏向尾部。在对称分布中,均值和中位数都位于靠近分布最高点的中心位置。

Discuss the mean, median, and mode for each of the following problems. Is there a pattern between the shape and measure of the center?

针对下面各题,讨论其均值、中位数和众数。形态与中心度量之间是否存在某种规律?

a\.

a.

b\.

b.

| The Ages Former U.S Presidents Died | |

| 美国前总统去世年龄 | |

|-------------------------------------|-------------------------|

|-------------------------------------|-------------------------|

| 4 | 6 9 |

| 4 | 6 9 |

| 5 | 3 6 7 7 7 8 |

| 5 | 3 6 7 7 7 8 |

| 6 | 0 0 3 3 4 4 5 6 7 7 7 8 |

| 6 | 0 0 3 3 4 4 5 6 7 7 7 8 |

| 7 | 0 1 1 2 3 4 7 8 8 9 |

| 7 | 0 1 1 2 3 4 7 8 8 9 |

| 8 | 0 1 3 5 8 |

| 8 | 0 1 3 5 8 |

| 9 | 0 0 3 3 |

| 9 | 0 0 3 3 |

| Key: 8\|0 means 80. | |

| 图例:8\|0 表示 80。 | |

Table 2.28

表 2.28

c\.

c.

2.7 Measures of the Spread of the Data 2.7 数据的离散程度度量

An important characteristic of any set of data is the variation in the data. In some data sets, the data values are concentrated closely near the mean; in other data sets, the data values are more widely spread out from the mean. The most common measure of variation, or spread, is the standard deviation. The standard deviation is a number that measures how far data values are from their mean.

任何数据集的一个重要特征都是数据的变异程度。在某些数据集中,数据值紧密聚集在均值附近;而在另一些数据集中,数据值则更分散地分布在均值之外。最常用的变异(或离散)度量是标准差。标准差是一个衡量数据值离其均值有多远的数字。

The standard deviation 标准差

The standard deviation provides a measure of the overall variation in a data set 标准差提供了对数据集中总体变异程度的度量

The standard deviation is always positive or zero. The standard deviation is small when the data are all concentrated close to the mean, exhibiting little variation or spread. The standard deviation is larger when the data values are more spread out from the mean, exhibiting more variation.

标准差总是为正或为零。当数据全部紧密聚集在均值附近时,标准差很小,表现出很小的变异或离散;当数据值更分散地分布在均值之外时,标准差较大,表现出更大的变异。

Suppose that we are studying the amount of time customers wait in line at the checkout at supermarket *A* and supermarket *B*. the average wait time at both supermarkets is five minutes. At supermarket *A*, the standard deviation for the wait time is two minutes; at supermarket *B* the standard deviation for the wait time is four minutes.

假设我们在研究顾客在超市 A 和超市 B 收银台排队等候的时间。两家超市的平均等候时间都是 5 分钟。在超市 A,等候时间的标准差为 2 分钟;在超市 B,等候时间的标准差为 4 分钟。

Because supermarket *B* has a higher standard deviation, we know that there is more variation in the wait times at supermarket *B*. Overall, wait times at supermarket *B* are more spread out from the average; wait times at supermarket *A* are more concentrated near the average.

因为超市 B 的标准差更大,我们知道超市 B 的等候时间变异更大。总体而言,超市 B 的等候时间更分散地偏离平均值;而超市 A 的等候时间更集中地分布在平均值附近。

The standard deviation can be used to determine whether a data value is close to or far from the mean. 标准差可用于判断某个数据值是接近还是远离均值。

Suppose that Rosa and Binh both shop at supermarket *A*. Rosa waits at the checkout counter for seven minutes and Binh waits for one minute. At supermarket *A*, the mean waiting time is five minutes and the standard deviation is two minutes. The standard deviation can be used to determine whether a data value is close to or far from the mean.

假设 Rosa 和 Binh 都在超市 A 购物。Rosa 在收银台等了 7 分钟,Binh 等了 1 分钟。在超市 A,平均等候时间为 5 分钟,标准差为 2 分钟。标准差可用于判断某个数据值是接近还是远离均值。

Rosa waits for seven minutes:

Rosa 等了 7 分钟:

Binh waits for one minute.

Binh 等了 1 分钟。

The number line may help you understand standard deviation. If we were to put five and seven on a number line, seven is to the right of five. We say, then, that seven is one standard deviation to the right of five because 5 + (1)(2) = 7.

数轴有助于你理解标准差。如果我们在数轴上标出 5 和 7,那么 7 在 5 的右边。于是我们说,7 在 5 的右边 1 个标准差处,因为 5 + (1)(2) = 7。

If one were also part of the data set, then one is two standard deviations to the left of five because 5 + (–2)(2) = 1.

如果 1 也是该数据集的一部分,那么 1 在 5 的左边 2 个标准差处,因为 5 + (–2)(2) = 1。

The equation value = mean + (#ofSTDEVs)(standard deviation) can be expressed for a sample and for a population.

等式 数值 = 均值 +(标准差个数)(标准差)可以分别针对样本和总体来表达。

The lower case letter *s* represents the sample standard deviation and the Greek letter *σ* (sigma, lower case) represents the population standard deviation.

小写字母 s 表示样本标准差,希腊字母 σ(sigma,小写)表示总体标准差。

The symbol $\overline{x}$ is the sample mean and the Greek symbol $\mu$ is the population mean.

符号 $\overline{x}$ 是样本均值,希腊符号 $\mu$ 是总体均值。

Calculating the Standard Deviation 计算标准差

If *x* is a number, then the difference "*x* – mean" is called its deviation. In a data set, there are as many deviations as there are items in the data set. The deviations are used to calculate the standard deviation. If the numbers belong to a population, in symbols a deviation is *x* – *μ*. For sample data, in symbols a deviation is *x* – $\overline{x}$.

如果 x 是一个数值,那么"x – 均值"的差称为它的偏差。在一个数据集中,偏差的个数与数据集中的项数一样多。这些偏差被用来计算标准差。如果这些数属于一个总体,则用符号表示偏差为 x – μ;对于样本数据,用符号表示偏差为 x – $\overline{x}$。

The procedure to calculate the standard deviation depends on whether the numbers are the entire population or are data from a sample. The calculations are similar, but not identical. Therefore the symbol used to represent the standard deviation depends on whether it is calculated from a population or a sample. The lower case letter s represents the sample standard deviation and the Greek letter *σ* (sigma, lower case) represents the population standard deviation. If the sample has the same characteristics as the population, then s should be a good estimate of *σ*.

计算标准差的步骤取决于这些数是整个总体还是来自样本的数据。两种计算相似但并不完全相同。因此,表示标准差的符号取决于它是根据总体还是样本计算得到的。小写字母 s 表示样本标准差,希腊字母 σ(sigma,小写)表示总体标准差。如果样本具有与总体相同的特征,那么 s 应当是对 σ 的一个良好估计。

To calculate the standard deviation, we need to calculate the variance first. The variance is the average of the squares of the deviations (the *x* – $\overline{x}$ values for a sample, or the *x* – *μ* values for a population). The symbol *σ*2 represents the population variance; the population standard deviation *σ* is the square root of the population variance. The symbol *s*2 represents the sample variance; the sample standard deviation *s* is the square root of the sample variance. You can think of the standard deviation as a special average of the deviations.

要计算标准差,我们需要先计算方差。方差是偏差平方的平均数(对样本为 x – $\overline{x}$ 的值,对总体为 x – μ 的值)。符号 σ2 表示总体方差;总体标准差 σ 是总体方差的平方根。符号 s2 表示样本方差;样本标准差 s 是样本方差的平方根。你可以把标准差看作偏差的一种特殊平均数。

If the numbers come from a census of the entire population and not a sample, when we calculate the average of the squared deviations to find the variance, we divide by *N*, the number of items in the population. If the data are from a sample rather than a population, when we calculate the average of the squared deviations, we divide by ***n* – 1**, one less than the number of items in the sample.

如果这些数据来自对整个总体的普查而非样本,那么当我们计算偏差平方的平均数以求方差时,除以 N(总体中的项数)。如果数据来自样本而非总体,那么当我们计算偏差平方的平均数时,除以 n – 1,即样本项数减 1。

Formulas for the Sample Standard Deviation 样本标准差公式

Formulas for the Population Standard Deviation 总体标准差公式

In these formulas, *f* represents the frequency with which a value appears. For example, if a value appears once, *f* is one. If a value appears three times in the data set or population, *f* is three.

在这些公式中,f 表示一个数值出现的频数。例如,如果某个值出现一次,f 为 1;如果某个值在该数据集或总体中出现三次,f 为 3。

Sampling Variability of a Statistic 统计量的抽样变异性

The statistic of a sampling distribution was discussed in Descriptive Statistics: Measuring the Center of the Data. How much the statistic varies from one sample to another is known as the sampling variability of a statistic. You typically measure the sampling variability of a statistic by its standard error. The standard error of the mean is an example of a standard error. It is a special standard deviation and is known as the standard deviation of the sampling distribution of the mean. You will cover the standard error of the mean in the chapter The Central Limit Theorem (not now). The notation for the standard error of the mean is $\frac{\sigma}{\sqrt{n}}$ where *σ* is the standard deviation of the population and n is the size of the sample.

抽样分布的统计量已在《描述统计学:度量数据的中心》中讨论过。统计量在不同样本之间变化的程度称为统计量的抽样变异性。你通常用一个统计量的标准误来度量其抽样变异性。均值的标准误就是标准误的一个例子。它是一种特殊的标准差,称为均值抽样分布的标准差。你将在《中心极限定理》一章(不是现在)中学习均值的标准误。均值标准误的记号为 $\frac{\sigma}{\sqrt{n}}$,其中 *σ* 为总体标准差,n 为样本量。

**In practice, USE A CALCULATOR OR COMPUTER SOFTWARE TO CALCULATE THE STANDARD DEVIATION. If you are using a TI-83, 83+, 84+ calculator, you need to select the appropriate standard deviation *σx* or *sx* from the summary statistics.** We will concentrate on using and interpreting the information that the standard deviation gives us. However you should study the following step-by-step example to help you understand how the standard deviation measures variation from the mean. (The calculator instructions appear at the end of this example.)

**在实践中,使用计算器或计算机软件来计算标准差。如果你使用的是 TI-83、83+、84+ 计算器,需要从汇总统计量中选择合适的标准差 *σx* 或 *sx*。** 我们将重点放在使用和解释标准差提供给我们的信息上。不过,你应该学习下面这个循序渐进的例子,以帮助你理解标准差是如何度量数据相对于均值的变异程度的。(计算器的操作说明出现在本例的末尾。)

In a fifth grade class, the teacher was interested in the average age and the sample standard deviation of the ages of her students. The following data are the ages for a SAMPLE of *n* = 20 fifth grade students. The ages are rounded to the nearest half year:

在一个五年级班级里,老师对她学生年龄的平均值和样本标准差感兴趣。以下数据是一个由 *n* = 20 名五年级学生组成的样本的年龄。年龄已四舍五入到最近的半岁:

9; 9.5; 9.5; 10; 10; 10; 10; 10.5; 10.5; 10.5; 10.5; 11; 11; 11; 11; 11; 11; 11.5; 11.5; 11.5;

9;9.5;9.5;10;10;10;10;10.5;10.5;10.5;10.5;11;11;11;11;11;11;11.5;11.5;11.5;

$$\overline{x} = \frac{\text{9~+~9}\text{.5(2)~+~10(4)~+~10}\text{.5(4)~+~11(6)~+~11}\text{.5(3)}}{20} = 10.525$$

样本均值(算术平均)为 $$\overline{x} = \frac{\text{9~+~9}\text{.5(2)~+~10(4)~+~10}\text{.5(4)~+~11(6)~+~11}\text{.5(3)}}{20} = 10.525$$

The average age is 10.53 years, rounded to two places.

平均年龄是 10.53 岁,四舍五入到两位小数。

The variance may be calculated by using a table. Then the standard deviation is calculated by taking the square root of the variance. We will explain the parts of the table after calculating *s*.

方差可以用一个表格来计算。然后取方差的平方根得到标准差。我们将在计算完 *s* 之后解释表格的各个部分。

| Data | Freq. | Deviations | *Deviations*2 | (Freq.)(*Deviations*2) |

| 数据 | 频数 | 离差 | *离差*2 | (频数)(*离差*2) |

|------|-------|------------------------|------------------------------------|-----------------------------------------|

|------|-------|------------------------|------------------------------------|-----------------------------------------|

| *x* | *f* | (*x* – $\overline{x}$) | (*x* – $\overline{x}$)2 | (*f*)(*x* – $\overline{x}$)2 |

| *x* | *f* | (*x* – $\overline{x}$) | (*x* – $\overline{x}$)2 | (*f*)(*x* – $\overline{x}$)2 |

| 9 | 1 | 9 – 10.525 = –1.525 | (–1.525)2 = 2.325625 | 1 × 2.325625 = 2.325625 |

| 9 | 1 | 9 – 10.525 = –1.525 | (–1.525)2 = 2.325625 | 1 × 2.325625 = 2.325625 |

| 9.5 | 2 | 9.5 – 10.525 = –1.025 | (–1.025)2 = 1.050625 | 2 × 1.050625 = 2.101250 |

| 9.5 | 2 | 9.5 – 10.525 = –1.025 | (–1.025)2 = 1.050625 | 2 × 1.050625 = 2.101250 |

| 10 | 4 | 10 – 10.525 = –0.525 | (–0.525)2 = 0.275625 | 4 × 0.275625 = 1.1025 |

| 10 | 4 | 10 – 10.525 = –0.525 | (–0.525)2 = 0.275625 | 4 × 0.275625 = 1.1025 |

| 10.5 | 4 | 10.5 – 10.525 = –0.025 | (–0.025)2 = 0.000625 | 4 × 0.000625 = 0.0025 |

| 10.5 | 4 | 10.5 – 10.525 = –0.025 | (–0.025)2 = 0.000625 | 4 × 0.000625 = 0.0025 |

| 11 | 6 | 11 – 10.525 = 0.475 | (0.475)2 = 0.225625 | 6 × 0.225625 = 1.35375 |

| 11 | 6 | 11 – 10.525 = 0.475 | (0.475)2 = 0.225625 | 6 × 0.225625 = 1.35375 |

| 11.5 | 3 | 11.5 – 10.525 = 0.975 | (0.975)2 = 0.950625 | 3 × 0.950625 = 2.851875 |

| 11.5 | 3 | 11.5 – 10.525 = 0.975 | (0.975)2 = 0.950625 | 3 × 0.950625 = 2.851875 |

| | | | | The total is 9.7375 |

| | | | | 总计为 9.7375 |

Table 2.29

表 2.29

The sample variance, *s*2, is equal to the sum of the last column (9.7375) divided by the total number of data values minus one (20 – 1):

样本方差 *s*2 等于最后一列(9.7375)之和除以数据值总数减一(20 – 1):

$s^{2} = \frac{9.7375}{20 - 1} = 0.5125$

样本方差为 $s^{2} = \frac{9.7375}{20 - 1} = 0.5125$

The sample standard deviation *s* is equal to the square root of the sample variance:

样本标准差 *s* 等于样本方差的平方根:

$s = \sqrt{0.5125} = 0.715891,$ which is rounded to two decimal places, *s* = 0.72.

$s = \sqrt{0.5125} = 0.715891$(四舍五入到两位小数),*s* = 0.72。

Typically, you do the calculation for the standard deviation on your calculator or computer. The intermediate results are not rounded. This is done for accuracy.

通常,你在计算器或计算机上完成标准差的计算。中间结果不做四舍五入,这是为保证准确性。

Problem 问题

1. Verify the mean and standard deviation on your calculator or computer.

1. 用计算器或计算机验证均值和标准差。

2. Find the value that is one standard deviation above the mean. Find ($\overline{x}$ + 1s).

2. 求高于均值一个标准差的数值。求 ($\overline{x}$ + 1s)。

3. Find the value that is two standard deviations below the mean. Find ($\overline{x}$ – 2s).

3. 求低于均值两个标准差的数值。求 ($\overline{x}$ – 2s)。

4. Find the values that are 1.5 standard deviations from (below and above) the mean.

4. 求距离均值(下方和上方)1.5 个标准差的数值。

Solution 解答

1. - Clear lists L1 and L2. Press STAT 4:ClrList. Enter 2nd 1 for L1, the comma (,), and 2nd 2 for L2.

1. - 清除列表 L1 和 L2。按 STAT 4:ClrList。输入 2nd 1 代表 L1,逗号(,),以及 2nd 2 代表 L2。

2. ($\overline{x}$ + 1s) = 10.53 + (1)(0.72) = 11.25

2. ($\overline{x}$ + 1s) = 10.53 + (1)(0.72) = 11.25

3. ($\overline{x}$ – 2*s*) = 10.53 – (2)(0.72) = 9.09

3. ($\overline{x}$ – 2*s*) = 10.53 – (2)(0.72) = 9.09

4. - ($\overline{x}$ – 1.5*s*) = 10.53 – (1.5)(0.72) = 9.45

4. - ($\overline{x}$ – 1.5*s*) = 10.53 – (1.5)(0.72) = 9.45

On a baseball team, the ages of each of the players are as follows:

在一支棒球队中,每名队员的年龄如下:

21; 21; 22; 23; 24; 24; 25; 25; 28; 29; 29; 31; 32; 33; 33; 34; 35; 36; 36; 36; 36; 38; 38; 38; 40

21;21;22;23;24;24;25;25;28;29;29;31;32;33;33;34;35;36;36;36;36;38;38;38;40

Use your calculator or computer to find the mean and standard deviation. Then find the value that is two standard deviations above the mean.

用你的计算器或计算机求出均值和标准差。然后求出高于均值两个标准差的数值。

Explanation of the standard deviation calculation shown in the table 表中标准差计算方法的解释

The deviations show how spread out the data are about the mean. The data value 11.5 is farther from the mean than is the data value 11 which is indicated by the deviations 0.97 and 0.47. A positive deviation occurs when the data value is greater than the mean, whereas a negative deviation occurs when the data value is less than the mean. The deviation is –1.525 for the data value nine. If you add the deviations, the sum is always zero. (For Example 2.32, there are *n* = 20 deviations.) So you cannot simply add the deviations to get the spread of the data. By squaring the deviations, you make them positive numbers, and the sum will also be positive. The variance, then, is the average squared deviation.

离差反映出数据围绕均值的离散程度。数据值 11.5 比数据值 11 离均值更远,这一点由离差 0.97 和 0.47 体现出来。当数据值大于均值时出现正离差,而当数据值小于均值时出现负离差。数据值 9 的离差为 –1.525。若把离差相加,其和总是零。(对于例 2.32,有 *n* = 20 个离差。)因此你不能简单地把离差相加来得到数据的离散程度。将离差平方,就使它们成为正数,其和也为正数。于是,方差就是离差平方的平均。

The variance is a squared measure and does not have the same units as the data. Taking the square root solves the problem. The standard deviation measures the spread in the same units as the data.

方差是一种平方度量,其单位与数据不同。取平方根解决了这个问题。标准差以与数据相同的单位来度量离散程度。

Notice that instead of dividing by *n* = 20, the calculation divided by *n* – 1 = 20 – 1 = 19 because the data is a sample. For the sample variance, we divide by the sample size minus one (*n* – 1). Why not divide by *n*? The answer has to do with the population variance. The sample variance is an estimate of the population variance. Based on the theoretical mathematics that lies behind these calculations, dividing by (*n* – 1) gives a better estimate of the population variance.

注意,计算不是除以 *n* = 20,而是除以 *n* – 1 = 20 – 1 = 19,因为数据是一个样本。对于样本方差,我们除以样本量减一(*n* – 1)。为什么不除以 *n*?答案与总体方差有关。样本方差是总体方差的一个估计。基于这些计算背后的理论数学,除以(*n* – 1)能给出对总体方差更好的估计。

Your concentration should be on what the standard deviation tells us about the data. The standard deviation is a number which measures how far the data are spread from the mean. Let a calculator or computer do the arithmetic.

你的注意力应放在标准差告诉我们关于数据的什么信息上。标准差是一个度量数据相对于均值离散程度的数值。让计算器或计算机来完成算术运算。

The standard deviation, *s* or *σ*, is either zero or larger than zero. Describing the data with reference to the spread is called "variability". The variability in data depends upon the method by which the outcomes are obtained; for example, by measuring or by random sampling. When the standard deviation is zero, there is no spread; that is, the all the data values are equal to each other. The standard deviation is small when the data are all concentrated close to the mean, and is larger when the data values show more variation from the mean. When the standard deviation is a lot larger than zero, the data values are very spread out about the mean; outliers can make *s* or *σ* very large.

标准差 *s* 或 *σ* 要么为零,要么大于零。参照离散程度来描述数据称为"变异性"(variability)。数据的变异性取决于获得结果的方法,例如通过测量或随机抽样。当标准差为零时,没有离散程度,也就是说所有数据值彼此相等。当数据都集中在均值附近时,标准差较小;当数据值相对于均值表现出更大变异时,标准差较大。当标准差远大于零时,数据值围绕均值非常分散;离群值会使 *s* 或 *σ* 变得非常大。

The standard deviation, when first presented, can seem unclear. By graphing your data, you can get a better "feel" for the deviations and the standard deviation. You will find that in symmetrical distributions, the standard deviation can be very helpful but in skewed distributions, the standard deviation may not be much help. The reason is that the two sides of a skewed distribution have different spreads. In a skewed distribution, it is better to look at the first quartile, the median, the third quartile, the smallest value, and the largest value. Because numbers can be confusing, always graph your data. Display your data in a histogram or a box plot.

标准差在初次出现时可能显得不清楚。通过绘制数据的图形,你可以对离差和标准差有更好的"感觉"。你会发现,在对称分布中,标准差可能很有帮助,但在偏态分布中,标准差可能帮助不大。原因是偏态分布的两侧具有不同的离散程度。在偏态分布中,最好观察第一四分位数、中位数、第三四分位数、最小值和最大值。因为数字可能让人困惑,始终要将你的数据作图。用直方图或箱线图来展示你的数据。

Problem 问题

Use the following data (first exam scores) from Susan Dean's spring pre-calculus class:

使用下列来自 Susan Dean 春季预科微积分课的数据(第一次考试成绩):

33; 42; 49; 49; 53; 55; 55; 61; 63; 67; 68; 68; 69; 69; 72; 73; 74; 78; 80; 83; 88; 88; 88; 90; 92; 94; 94; 94; 94; 96; 100

33;42;49;49;53;55;55;61;63;67;68;68;69;69;72;73;74;78;80;83;88;88;88;90;92;94;94;94;94;96;100

1. Create a chart containing the data, frequencies, relative frequencies, and cumulative relative frequencies to three decimal places.

1. 创建一个包含数据、频数、相对频数和累积相对频数(保留三位小数)的图表。

2. Calculate the following to one decimal place using a TI-83+ or TI-84 calculator:

2. 用 TI-83+ 或 TI-84 计算器计算下列各项,保留一位小数:

1. The sample mean

1. 样本均值

2. The sample standard deviation

2. 样本标准差

3. The median

3. 中位数

4. The first quartile

4. 第一四分位数

5. The third quartile

5. 第三四分位数

6. *IQR*

6. *IQR*

3. Construct a box plot and a histogram on the same set of axes. Make comments about the box plot, the histogram, and the chart.

3. 在同一组坐标轴上绘制箱线图和直方图。对箱线图、直方图和图表作出评论。

Solution 解答

1. See Table 2.30

1. 见表 2.30

2. 1. The sample mean = 73.5

2. 1. 样本均值 = 73.5

2. The sample standard deviation = 17.9

2. 样本标准差 = 17.9

3. The median = 73

3. 中位数 = 73

4. The first quartile = 61

4. 第一四分位数 = 61

5. The third quartile = 90

5. 第三四分位数 = 90

6. *IQR* = 90 – 61 = 29

6. *IQR* = 90 – 61 = 29

3. The *x*-axis goes from 32.5 to 100.5; *y*-axis goes from –2.4 to 15 for the histogram. The number of intervals is five, so the width of an interval is (100.5 – 32.5) divided by five, is equal to 13.6. Endpoints of the intervals are as follows: the starting point is 32.5, 32.5 + 13.6 = 46.1, 46.1 + 13.6 = 59.7, 59.7 + 13.6 = 73.3, 73.3 + 13.6 = 86.9, 86.9 + 13.6 = 100.5 = the ending value; No data values fall on an interval boundary.

3. 直方图的水平轴从 32.5 到 100.5;垂直轴从 –2.4 到 15。区间数为 5,因此区间宽度为 (100.5 – 32.5) 除以 5,等于 13.6。各区间的端点如下:起点为 32.5,32.5 + 13.6 = 46.1,46.1 + 13.6 = 59.7,59.7 + 13.6 = 73.3,73.3 + 13.6 = 86.9,86.9 + 13.6 = 100.5 = 终点值;没有数据值落在区间边界上。

| Data | Frequency | Relative Frequency | Cumulative Relative Frequency |

| 数据 | 频数 | 相对频数 | 累积相对频数 |

|------|-----------|--------------------|---------------------------------|

|------|-----------|--------------------|---------------------------------|

| 33 | 1 | 0.032 | 0.032 |

| 33 | 1 | 0.032 | 0.032 |

| 42 | 1 | 0.032 | 0.064 |

| 42 | 1 | 0.032 | 0.064 |

| 49 | 2 | 0.065 | 0.129 |

| 49 | 2 | 0.065 | 0.129 |

| 53 | 1 | 0.032 | 0.161 |

| 53 | 1 | 0.032 | 0.161 |

| 55 | 2 | 0.065 | 0.226 |

| 55 | 2 | 0.065 | 0.226 |

| 61 | 1 | 0.032 | 0.258 |

| 61 | 1 | 0.032 | 0.258 |

| 63 | 1 | 0.032 | 0.29 |

| 63 | 1 | 0.032 | 0.29 |

| 67 | 1 | 0.032 | 0.322 |

| 67 | 1 | 0.032 | 0.322 |

| 68 | 2 | 0.065 | 0.387 |

| 68 | 2 | 0.065 | 0.387 |

| 69 | 2 | 0.065 | 0.452 |

| 69 | 2 | 0.065 | 0.452 |

| 72 | 1 | 0.032 | 0.484 |

| 72 | 1 | 0.032 | 0.484 |

| 73 | 1 | 0.032 | 0.516 |

| 73 | 1 | 0.032 | 0.516 |

| 74 | 1 | 0.032 | 0.548 |

| 74 | 1 | 0.032 | 0.548 |

| 78 | 1 | 0.032 | 0.580 |

| 78 | 1 | 0.032 | 0.580 |

| 80 | 1 | 0.032 | 0.612 |

| 80 | 1 | 0.032 | 0.612 |

| 83 | 1 | 0.032 | 0.644 |

| 83 | 1 | 0.032 | 0.644 |

| 88 | 3 | 0.097 | 0.741 |

| 88 | 3 | 0.097 | 0.741 |

| 90 | 1 | 0.032 | 0.773 |

| 90 | 1 | 0.032 | 0.773 |

| 92 | 1 | 0.032 | 0.805 |

| 92 | 1 | 0.032 | 0.805 |

| 94 | 4 | 0.129 | 0.934 |

| 94 | 4 | 0.129 | 0.934 |

| 96 | 1 | 0.032 | 0.966 |

| 96 | 1 | 0.032 | 0.966 |

| 100 | 1 | 0.032 | 0.998 (Why isn't this value 1?) |

| 100 | 1 | 0.032 | 0.998(为什么这个值不是 1?) |

Table 2.30

表 2.30

The long left whisker in the box plot is reflected in the left side of the histogram. The spread of the exam scores in the lower 50% is greater (73 – 33 = 40) than the spread in the upper 50% (100 – 73 = 27). The histogram, box plot, and chart all reflect this. There are a substantial number of A and B grades (80s, 90s, and 100). The histogram clearly shows this. The box plot shows us that the middle 50% of the exam scores (*IQR* = 29) are Ds, Cs, and Bs. The box plot also shows us that the lower 25% of the exam scores are Ds and Fs.

箱线图中较长的左须反映在直方图的左侧。考试成绩在下半部分 50% 的离散程度(73 – 33 = 40)大于上半部分 50%(100 – 73 = 27)。直方图、箱线图和图表都反映了这一点。有相当数量的 A 和 B 等级(80 多分、90 多分和 100 分)。直方图清楚地显示了这一点。箱线图告诉我们,考试成绩的中间 50%(*IQR* = 29)是 D、C 和 B。箱线图还告诉我们,考试成绩最低的 25% 是 D 和 F。

The following data show the different types of pet food stores in the area carry.

下列数据显示该地区宠物食品商店所售的不同种类。

6; 6; 6; 6; 7; 7; 7; 7; 7; 8; 9; 9; 9; 9; 10; 10; 10; 10; 10; 11; 11; 11; 11; 12; 12; 12; 12; 12; 12;

6;6;6;6;7;7;7;7;7;8;9;9;9;9;10;10;10;10;10;11;11;11;11;12;12;12;12;12;12;

Calculate the sample mean and the sample standard deviation to one decimal place using a TI-83+ or TI-84 calculator.

用 TI-83+ 或 TI-84 计算器计算样本均值和样本标准差,保留一位小数。

Standard deviation of Grouped Frequency Tables 分组频数表的标准差

Recall that for grouped data we do not know individual data values, so we cannot describe the typical value of the data with precision. In other words, we cannot find the exact mean, median, or mode. We can, however, determine the best estimate of the measures of center by finding the mean of the grouped data with the formula: $Mean\ of\ Frequency\ Table = \frac{\sum{fm}}{\sum f}$

回顾一下,对于分组数据,我们不知道各个数据值,因此无法精确地描述数据的典型值。换言之,我们无法求出准确的均值、中位数或众数。不过,我们可以通过求分组数据的均值来得出集中趋势度量的最佳估计,公式如下:$Mean\ of\ Frequency\ Table = \frac{\sum{fm}}{\sum f}$

where $f =$ interval frequencies and *m* = interval midpoints.

其中 $f =$ 组频数,*m* = 组中点。

Just as we could not find the exact mean, neither can we find the exact standard deviation. Remember that standard deviation describes numerically the expected deviation a data value has from the mean. In simple English, the standard deviation allows us to compare how "unusual" individual data is compared to the mean.

正如我们无法求出准确的均值一样,我们也无法求出准确的标准差。要记住,标准差是用数值描述某个数据值偏离均值的预期偏差。简单地说,标准差使我们能够比较单个数据与均值相比"不寻常"的程度。

Find the standard deviation for the data in Table 2.31.

求表 2.31 中数据的标准差。

| Class | Frequency, $f$ | Midpoint, $m$ | $f \cdot m$ | $\overline{x}$ | $m - \overline{x}$ | $\left( m - \overline{x} \right)^{2}$ | $f\left( m - \overline{x} \right)^{2}$ |

| 组 | 频数 $f$ | 中点 $m$ | $f \cdot m$ | $\overline{x}$ | $m - \overline{x}$ | $\left( m - \overline{x} \right)^{2}$ | $f\left( m - \overline{x} \right)^{2}$ |

|-----------------------------|----------------|---------------|-------------|----------------|--------------------|---------------------------------------|----------------------------------------|

|-----------------------------|----------------|---------------|-------------|----------------|--------------------|---------------------------------------|----------------------------------------|

| 0–2 | 1 | 1 | 1 | 7.58 | -6.58 | 43.2964 | 43.2964 |

| 0–2 | 1 | 1 | 1 | 7.58 | -6.58 | 43.2964 | 43.2964 |

| 3–5 | 6 | 4 | 24 | 7.58 | -3.58 | 12.8164 | 76.8984 |

| 3–5 | 6 | 4 | 24 | 7.58 | -3.58 | 12.8164 | 76.8984 |

| 6–8 | 10 | 7 | 70 | 7.58 | -0.58 | 0.3364 | 3.364 |

| 6–8 | 10 | 7 | 70 | 7.58 | -0.58 | 0.3364 | 3.364 |

| 9–11 | 7 | 10 | 70 | 7.58 | 2.42 | 5.8564 | 40.9948 |

| 9–11 | 7 | 10 | 70 | 7.58 | 2.42 | 5.8564 | 40.9948 |

| 12–14 | 0 | 13 | 0 | 7.58 | 5.42 | 29.3764 | 0 |

| 12–14 | 0 | 13 | 0 | 7.58 | 5.42 | 29.3764 | 0 |

| 15–17 | 2 | 16 | 32 | 7.58 | 8.42 | 70.8964 | 141.7928 |

| 15–17 | 2 | 16 | 32 | 7.58 | 8.42 | 70.8964 | 141.7928 |

| SUM ($\mathbf{\Sigma}$) | 26 | | 197 | | | | 43.2964 |

| 合计 ($\mathbf{\Sigma}$) | 26 | | 197 | | | | 43.2964 |

Table 2.31

表 2.31

The values in the second, third, and fourth columns of Table 2.31 are used to calculate the mean of the grouped frequency table, the value in the fifth column.

表 2.31 的第二、第三和第四列中的数值用于计算分组频数表的均值,即第五列中的数值。

$$\overline{x} = \frac{\Sigma fm}{\Sigma f} = \frac{197}{26} \approx 7.58.$$ 2.1

$$\overline{x} = \frac{\Sigma fm}{\Sigma f} = \frac{197}{26} \approx 7.58.$$ 2.1

After calculating $\overline{x}$, find the difference, $m - \overline{x}$, for each midpoint, $m$. Next, square each difference. In the final column, calculate the product of the frequency and the squared difference for each class.

计算出 $\overline{x}$ 后,对每个中点 $m$ 求差 $m - \overline{x}$。然后对每个差值平方。在最后一列中,计算各组频数与平方差的乘积。

The table makes it easy to use the formula for calculating the standard deviation of a grouped frequency table:

该表使得使用公式计算分组频数表的标准差变得容易:

$$s_{x} = \sqrt{\frac{\Sigma f\left( m - \overline{x} \right)^{2}}{n - 1}} = \sqrt{\frac{43.2964}{26 - 1}} \approx 3.50.$$ 2.2

$$s_{x} = \sqrt{\frac{\Sigma f\left( m - \overline{x} \right)^{2}}{n - 1}} = \sqrt{\frac{43.2964}{26 - 1}} \approx 3.50.$$ 2.2

Although the formula is not complicated, these calculations are typically performed using technology.

尽管该公式并不复杂,但这些计算通常使用技术工具(计算器或统计软件)完成。

Find the standard deviation for the data from the previous example

求上一例数据的标准差。

| Class | Frequency, *f* |

| 组 | 频数 *f* |

|-------|----------------|

|-------|----------------|

| 0–2 | 1 |

| 0–2 | 1 |

| 3–5 | 6 |

| 3–5 | 6 |

| 6–8 | 10 |

| 6–8 | 10 |

| 9–11 | 7 |

| 9–11 | 7 |

| 12–14 | 0 |

| 12–14 | 0 |

| 15–17 | 2 |

| 15–17 | 2 |

Table 2.32

表 2.32

First, press the STAT key and select 1:Edit

首先按 STAT(统计)键,选择 1:Edit(编辑)。

Input the midpoint values into L1 and the frequencies into L2

将中点数值输入 L1,将频数输入 L2

Select STAT, CALC, and 1: 1-Var Stats

选择 STAT(统计)、CALC(计算)和 1: 1-Var Stats(单变量统计)。

Select 2nd then 1 then , 2nd then 2 Enter

选择 2nd(第二功能)键,然后按 1,再按逗号,然后按 2nd 再按 2 Enter(回车)。

You will see displayed both a population standard deviation, *σx*, and the sample standard deviation, *sx*.

屏幕上将同时显示总体标准差 *σx* 和样本标准差 *sx*。

Comparing Values from Different Data Sets 比较来自不同数据集的数值

The standard deviation is useful when comparing data values that come from different data sets. If the data sets have different means and standard deviations, then comparing the data values directly can be misleading.

当比较来自不同数据集的数据值时,标准差很有用。如果各数据集的均值和标准差不同,那么直接比较数据值可能产生误导。

\#ofSTDEVs is often called a "*z*-score"; we can use the symbol *z*. In symbols, the formulas become:

\#ofSTDEVs 常被称为"*z* 分数"(*z*-score);我们可以用符号 *z* 表示。用符号表示,公式变为:

| | | |

| | | |

|------------|-----------------------------|--------------------------------------|

|------------|-----------------------------|--------------------------------------|

| Sample | $x$ = $\overline{x}$ + *zs* | $z = \frac{x\ - \ \overline{x}}{s}$ |

| 样本 | $x$ = $\overline{x}$ + *zs* | $z = \frac{x\ - \ \overline{x}}{s}$ |

| Population | $x$ = $\mu$ + *zσ* | $z = \frac{x\ - \ \mu}{\sigma}$ |

| 总体 | $x$ = $\mu$ + *zσ* | $z = \frac{x\ - \ \mu}{\sigma}$ |

Table 2.33

表 2.33

Problem 问题

Two students, John and Ali, from different high schools, wanted to find out who had the highest GPA when compared to his school. Which student had the highest GPA when compared to his school?

两名来自不同高中的学生 John 和 Ali 想要弄清楚,与各自学校相比,谁的 GPA 最高。与各自学校相比,哪名学生的 GPA 最高?

| Student | GPA | School Mean GPA | School Standard Deviation |

| 学生 | GPA | 学校平均 GPA | 学校标准差 |

|---------|------|-----------------|---------------------------|

|---------|------|-----------------|---------------------------|

| John | 2.85 | 3.0 | 0.7 |

| John | 2.85 | 3.0 | 0.7 |

| Ali | 77 | 80 | 10 |

| Ali | 77 | 80 | 10 |

Table 2.34

表 2.34

Solution 解答

For each student, determine how many standard deviations (#ofSTDEVs) his GPA is away from the average, for his school. Pay careful attention to signs when comparing and interpreting the answer.

对每名学生,确定其 GPA 偏离所在学校平均值多少个标准差(#ofSTDEVs)。在比较和解释答案时,要特别注意符号。

$z = \operatorname{\#\ of\ STDEVs} = \frac{\text{value~}–\text{mean}}{\text{standard~deviation}} = \frac{x–\mu}{\sigma}$

$z = \operatorname{\#\ of\ STDEVs} = \frac{\text{value~}–\text{mean}}{\text{standard~deviation}} = \frac{x–\mu}{\sigma}$

For John, $z = \# ofSTDEVs = \frac{2.85–3.0}{0.7} = –0.21$

For John, $z = \# ofSTDEVs = \frac{2.85–3.0}{0.7} = –0.21$

For Ali, $z = \# ofSTDEVs = \frac{77 - 80}{10} = - 0.3$

For Ali, $z = \# ofSTDEVs = \frac{77 - 80}{10} = - 0.3$

John has the better GPA when compared to his school because his GPA is 0.21 standard deviations below his school's mean while Ali's GPA is 0.3 standard deviations below his school's mean.

与各自学校相比,John 的 GPA 更好,因为他的 GPA 比学校均值低 0.21 个标准差,而 Ali 的 GPA 比学校均值低 0.3 个标准差。

John's *z*-score of –0.21 is higher than Ali's *z*-score of –0.3. For GPA, higher values are better, so we conclude that John has the better GPA when compared to his school.

John 的 *z* 分数 –0.21 高于 Ali 的 *z* 分数 –0.3。对于 GPA,数值越高越好,因此我们得出结论:与各自学校相比,John 的 GPA 更好。

Two swimmers, Angie and Beth, from different teams, wanted to find out who had the fastest time for the 50 meter freestyle when compared to her team. Which swimmer had the fastest time when compared to her team?

来自不同队伍的两名游泳运动员 Angie 和 Beth 想要弄清楚,与各自队伍相比,谁在 50 米自由泳中的成绩最快。与各自队伍相比,哪名游泳运动员的成绩最快?

| Swimmer | Time (seconds) | Team Mean Time | Team Standard Deviation |

| 游泳运动员 | 时间(秒) | 队伍平均时间 | 队伍标准差 |

|---------|----------------|----------------|-------------------------|

|---------|----------------|----------------|-------------------------|

| Angie | 26.2 | 27.2 | 0.8 |

| Angie | 26.2 | 27.2 | 0.8 |

| Beth | 27.3 | 30.1 | 1.4 |

| Beth | 27.3 | 30.1 | 1.4 |

Table 2.35

表 2.35

The following lists give a few facts that provide a little more insight into what the standard deviation tells us about the distribution of the data.

下面两份清单给出了一些事实,有助于进一步理解标准差揭示了数据分布的哪些信息。

For ANY data set, no matter what the distribution of the data is:

对于任意数据集,无论其数据分布如何:

For data having a distribution that is BELL-SHAPED and SYMMETRIC:

对于呈钟形且对称分布的数据:

2.8 Descriptive Statistics 2.8 描述统计学

Descriptive Statistics 描述统计学

Class Time:

上课时间:

Names:

姓名:

Student Learning Outcomes

学生学习目标

Collect the Data Record the number of pairs of shoes you own.

收集数据 记录你拥有的鞋子的双数。

1. Randomly survey 30 classmates about the number of pairs of shoes they own. Record their values.

1. 随机调查 30 名同学,了解他们拥有的鞋子双数,并记录这些数值。

| | | | | |

| | | | | |

|-----|-----|-----|-----|-----|

|-----|-----|-----|-----|-----|

| | | | | |

| | | | | |

| | | | | |

| | | | | |

| | | | | |

| | | | | |

| | | | | |

| | | | | |

| | | | | |

| | | | | |

| | | | | |

| | | | | |

Table 2.36 Survey Results

表 2.36 调查结果

2. Construct a histogram. Make five to six intervals. Sketch the graph using a ruler and pencil and scale the axes.

2. 绘制直方图。划分五到六个区间。用直尺和铅笔画出图形,并为坐标轴标上刻度。

3. Calculate the following values.

3. 计算下列数值。

1. $\overline{x}$ = \_\_\_\_\_

1. $\overline{x}$ = \_\_\_\_\_

2. *s* = \_\_\_\_\_

2. *s* = \_\_\_\_\_

4. Are the data discrete or continuous? How do you know?

4. 这些数据是离散的还是连续的?你是如何判断的?

5. In complete sentences, describe the shape of the histogram.

5. 用完整的句子描述直方图的形状。

6. Are there any potential outliers? List the value(s) that could be outliers. Use a formula to check the end values to determine if they are potential outliers.

6. 是否存在潜在的离群值?列出可能为离群值的数据。使用公式检验两端数值,以判断它们是否为潜在离群值。

Analyze the Data

分析数据

1. Determine the following values.

1. 确定下列数值。

1. Min = \_\_\_\_\_

1. Min = \_\_\_\_\_

2. *M* = \_\_\_\_\_

2. *M* = \_\_\_\_\_

3. Max = \_\_\_\_\_

3. Max = \_\_\_\_\_

4. *Q*1 = \_\_\_\_\_

4. *Q*1 = \_\_\_\_\_

5. *Q*3 = \_\_\_\_\_

5. *Q*3 = \_\_\_\_\_

6. *IQR* = \_\_\_\_\_

6. *IQR* = \_\_\_\_\_

2. Construct a box plot of data

2. 绘制数据的箱线图

3. What does the shape of the box plot imply about the concentration of data? Use complete sentences.

3. 箱线图的形状暗示了数据的集中情况如何?请使用完整的句子作答。

4. Using the box plot, how can you determine if there are potential outliers?

4. 利用箱线图,你如何判断是否存在潜在的离群值?

5. How does the standard deviation help you to determine concentration of the data and whether or not there are potential outliers?

5. 标准差如何帮助你确定数据的集中程度以及是否存在潜在的离群值?

6. What does the *IQR* represent in this problem?

6. 在此问题中,*IQR* 代表什么?

7. Show your work to find the value that is 1.5 standard deviations:

7. 展示你的计算过程,求下列数值(距均值 1.5 个标准差):

1. above the mean.

1. 高于均值。

2. below the mean.

2. 低于均值。

Key Terms 关键术语

Box plot

箱线图

a graph that gives a quick picture of the middle 50% of the data

一种能快速呈现数据中间 50% 分布状况的图形。

First Quartile

第一四分位数

the value that is the median of the of the lower half of the ordered data set

有序数据集下半部分的中位数。

Frequency

频数

the number of times a value of the data occurs

某个数据值出现的次数。

Frequency Polygon

频数多边形

looks like a line graph but uses intervals to display ranges of large amounts of data

形似线图,但使用区间来展示大量数据的范围。

Frequency Table

频数表

a data representation in which grouped data is displayed along with the corresponding frequencies

一种数据表示方式,将分组数据连同相应的频数一起展示。

Histogram

直方图

a graphical representation in *x*-*y* form of the distribution of data in a data set; *x* represents the data and *y* represents the frequency, or relative frequency. The graph consists of contiguous rectangles.

以 *x*-*y* 形式表示数据集中数据分布的一种图形;*x* 表示数据,*y* 表示频数或相对频数。该图形由相邻的矩形组成。

Interquartile Range

四分位距

or *IQR*, is the range of the middle 50 percent of the data values; the *IQR* is found by subtracting the first quartile from the third quartile.

又称 *IQR*,是数据中间 50% 的数值的范围;*IQR* 由第三四分位数减去第一四分位数得到。

Interval

区间

also called a class interval; an interval represents a range of data and is used when displaying large data sets

也称组区间(class interval);区间表示一段数据范围,在展示大型数据集时使用。

Mean

均值

a number that measures the central tendency of the data; a common name for mean is 'average.' The term 'mean' is a shortened form of 'arithmetic mean.' By definition, the mean for a sample (denoted by $\overline{x}$) is $\overline{x}\ = \ \frac{\text{Sum~of~all~values~in~the~sample}}{\text{Number~of~values~in~the~sample}}$, and the mean for a population (denoted by *μ*) is $\mu = \frac{\text{Sum~of~all~values~in~the~population}}{\text{Number~of~values~in~the~population}}$.

衡量数据集中趋势的一个数值;均值常用的名称是"平均数"(average)。"mean"是"arithmetic mean"(算术平均)的简写。根据定义,样本均值(记为 $\overline{x}$)为 $\overline{x}\ = \ \frac{\text{Sum~of~all~values~in~the~sample}}{\text{Number~of~values~in~the~sample}}$,总体均值(记为 *μ*)为 $\mu = \frac{\text{Sum~of~all~values~in~the~population}}{\text{Number~of~values~in~the~population}}$。

Median

中位数

a number that separates ordered data into halves; half the values are the same number or smaller than the median and half the values are the same number or larger than the median. The median may or may not be part of the data.

将有序数据分成两半的一个数值;一半的数值与该数相等或小于中位数,另一半与该数相等或大于中位数。中位数可能是也可能不是数据本身的一部分。

Midpoint

中点

the mean of an interval in a frequency table

频数表中某个区间的均值。

Mode

众数

the value that appears most frequently in a set of data

在一组数据中出现的次数最多的值。

Outlier

离群值

an observation that does not fit the rest of the data

与数据其余部分不相符的一个观测值。

Paired Data Set

配对数据集

two data sets that have a one to one relationship so that:

具有一一对应关系的两组数据,满足:

Percentile

百分位数

a number that divides ordered data into hundredths; percentiles may or may not be part of the data. The median of the data is the second quartile and the 50th percentile. The first and third quartiles are the 25th and the 75th percentiles, respectively.

将有序数据分成百分位的一个数值;百分位数可能是也可能不是数据本身的一部分。数据的中位数即第二四分位数,也是第 50th 百分位数。第一、第三四分位数分别是第 25th 和第 75th 百分位数。

Quartiles

四分位数

the numbers that separate the data into quarters; quartiles may or may not be part of the data. The second quartile is the median of the data.

将数据分成四等份的数值;四分位数可能是也可能不是数据本身的一部分。第二四分位数是数据的中位数。

Relative Frequency

相对频数

the ratio of the number of times a value of the data occurs in the set of all outcomes to the number of all outcomes

某个数据值在所有结果集合中出现的次数与所有结果总数之比。

Skewed

偏态的

used to describe data that is not symmetrical; when the right side of a graph looks "chopped off" compared the left side, we say it is "skewed to the left." When the left side of the graph looks "chopped off" compared to the right side, we say the data is "skewed to the right." Alternatively: when the lower values of the data are more spread out, we say the data are skewed to the left. When the greater values are more spread out, the data are skewed to the right.

用于描述非对称的数据;当图形的右侧相对于左侧看起来"被截去"时,称其"左偏"(skewed to the left)。当图形的左侧相对于右侧看起来"被截去"时,称数据"右偏"(skewed to the right)。另一种说法:当数据的较小值分布得更分散时,称数据左偏;当较大值分布得更分散时,数据右偏。

Standard Deviation

标准差

a number that is equal to the square root of the variance and measures how far data values are from their mean; notation: *s* for sample standard deviation and σ for population standard deviation.

等于方差平方根的一个数值,用于衡量数据值偏离其均值的程度;记号:*s* 表示样本标准差,σ 表示总体标准差。

Variance

方差

mean of the squared deviations from the mean, or the square of the standard deviation; for a set of data, a deviation can be represented as *x* – $\overline{x}$ where *x* is a value of the data and $\overline{x}$ is the sample mean. The sample variance is equal to the sum of the squares of the deviations divided by the difference of the sample size and one.

偏离均值的平方偏差的均值,即标准差的平方;对一组数据而言,偏差可表示为 *x* – $\overline{x}$,其中 *x* 是数据的一个值,$\overline{x}$ 是样本均值。样本方差等于偏差平方和除以样本量减一的差。

Chapter Review 章节回顾

2.1 Stem-and-Leaf Graphs (Stemplots), Line Graphs, and Bar Graphs 2.1 茎叶图(Stemplots)、线图与条形图

A stem-and-leaf plot is a way to plot data and look at the distribution. In a stem-and-leaf plot, all data values within a class are visible. The advantage in a stem-and-leaf plot is that all values are listed, unlike a histogram, which gives classes of data values. A line graph is often used to represent a set of data values in which a quantity varies with time. These graphs are useful for finding trends. That is, finding a general pattern in data sets including temperature, sales, employment, company profit or cost over a period of time. A bar graph is a chart that uses either horizontal or vertical bars to show comparisons among categories. One axis of the chart shows the specific categories being compared, and the other axis represents a discrete value. Some bar graphs present bars clustered in groups of more than one (grouped bar graphs), and others show the bars divided into subparts to show cumulative effect (stacked bar graphs). Bar graphs are especially useful when categorical data is being used.

茎叶图是一种绘制数据并观察其分布的方法。在茎叶图中,一个组(类)内的所有数据值都清晰可见。茎叶图的优点在于所有数值都被一一列出,而直方图只给出数据的各个组(类)。线图常用于表示一组随时间变化的数据值。这类图形有助于发现趋势,即在包括温度、销售额、就业、公司利润或成本等随时间变化的数据集中寻找总体模式。条形图是一种使用水平或垂直条形来显示各类别之间比较的图表。图表的一个轴显示被比较的具体类别,另一个轴表示离散数值。有些条形图将条形聚为多个一组(分组条形图),另一些则将条形分成若干子部分以显示累积效应(堆积条形图)。当使用分类数据时,条形图尤为有用。

2.2 Histograms, Frequency Polygons, and Time Series Graphs 2.2 直方图、频数多边形与时间序列图

A histogram is a graphic version of a frequency distribution. The graph consists of bars of equal width drawn adjacent to each other. The horizontal scale represents classes of quantitative data values and the vertical scale represents frequencies. The heights of the bars correspond to frequency values. Histograms are typically used for large, continuous, quantitative data sets. A frequency polygon can also be used when graphing large data sets with data points that repeat. The data usually goes on *y*-axis with the frequency being graphed on the *x*-axis. Time series graphs can be helpful when looking at large amounts of data for one variable over a period of time.

直方图是频数分布的一种图形化表示。该图形由等宽且相互邻接的条形组成。水平刻度表示定量数据值的各个组,垂直刻度表示频数。条形的高度对应于频数。直方图通常用于大型、连续、定量的数据集。当绘制含有重复数据点的大型数据集时,也可使用频数多边形。数据通常置于 *y* 轴,频数绘制在 *x* 轴。在考察某一变量随时间变化的大量数据时,时间序列图很有帮助。

2.3 Measures of the Location of the Data 2.3 数据的位置度量

The values that divide a rank-ordered set of data into 100 equal parts are called percentiles. Percentiles are used to compare and interpret data. For example, an observation at the 50th percentile would be greater than 50 percent of the other observations in the set. Quartiles divide data into quarters. The first quartile (*Q*1) is the 25th percentile,the second quartile (*Q*2 or median) is 50th percentile, and the third quartile (*Q*3) is the 75th percentile. The interquartile range, or *IQR*, is the range of the middle 50 percent of the data values. The *IQR* is found by subtracting *Q*1 from *Q*3, and can help determine outliers by using the following two expressions.

把一个按秩排序的数据集分成 100 个相等部分的值称为百分位数。百分位数用于比较和解释数据。例如,处于第 50 百分位数的观测值会大于该集合中 50% 的其他观测值。四分位数把数据分成四等份。第一四分位数(*Q*₁)是第 25 百分位数,第二四分位数(*Q*₂ 或中位数)是第 50 百分位数,第三四分位数(*Q*₃)是第 75 百分位数。四分位距(*IQR*)是中间 50% 数据值的范围。*IQR* 通过用 *Q*₃ 减去 *Q*₁ 得到,并可用下列两个表达式来帮助判断离群值。

2.4 Box Plots 2.4 箱线图

Box plots are a type of graph that can help visually organize data. To graph a box plot the following data points must be calculated: the minimum value, the first quartile, the median, the third quartile, and the maximum value. Once the box plot is graphed, you can display and compare distributions of data.

箱线图是一种能帮助直观地组织数据的图形。要绘制箱线图,必须先计算以下数据点:最小值、第一四分位数、中位数、第三四分位数和最大值。箱线图绘制完成后,便可展示并比较数据的分布。

2.5 Measures of the Center of the Data 2.5 数据中心位置的度量

The mean and the median can be calculated to help you find the "center" of a data set. The mean is the best estimate for the actual data set, but the median is the best measurement when a data set contains several outliers or extreme values. The mode will tell you the most frequently occurring datum (or data) in your data set. The mean, median, and mode are extremely helpful when you need to analyze your data, but if your data set consists of ranges which lack specific values, the mean may seem impossible to calculate. However, the mean can be approximated if you add the lower boundary with the upper boundary and divide by two to find the midpoint of each interval. Multiply each midpoint by the number of values found in the corresponding range. Divide the sum of these values by the total number of data values in the set.

可以计算均值和中位数来帮助你找到数据集的"中心"。均值是实际数据集的最佳估计,但当数据集中含有若干离群值或极端值时,中位数是最佳度量。众数会告诉你数据集中出现最频繁的数据点(datum 或 data)。均值、中位数和众数在你分析数据时极为有用;但如果你的数据集由缺乏具体数值的区间构成,均值似乎无法计算。不过,如果你把下边界与上边界相加再除以 2 以求得每个区间的中点,则均值可以近似得到。将每个中点乘以相应区间内数值的个数,再把这些数值之和除以数据值的总数。

2.6 Skewness and the Mean, Median, and Mode 2.6 偏度与均值、中位数、众数

Looking at the distribution of data can reveal a lot about the relationship between the mean, the median, and the mode. There are three types of distributions. A left (or negative) skewed distribution has a shape like Figure 2.17. A right (or positive) skewed distribution has a shape like Figure 2.18. A symmetrical distribution looks like Figure 2.16.

观察数据的分布可以揭示均值、中位数和众数之间关系的许多信息。分布共有三种类型。左偏(或负偏)分布的形状如图 2.17 所示。右偏(或正偏)分布的形状如图 2.18 所示。对称分布的形状如图 2.16 所示。

2.7 Measures of the Spread of the Data 2.7 数据离散程度的度量

The standard deviation can help you calculate the spread of data. There are different equations to use if are calculating the standard deviation of a sample or of a population.

标准差可以帮助你计算数据的离散程度。在计算样本的标准差或总体的标准差时,所使用的公式有所不同。

Formula Review 公式回顾

2.3 Measures of the Location of the Data 2.3 数据的位置度量

$i = \left( \frac{k}{100} \right)\left( {n + 1} \right)$

$i = \left( \frac{k}{100} \right)\left( {n + 1} \right)$

where *i* = the ranking or position of a data value,

其中 *i* = 数据值的秩或位置,

*k* = the kth percentile,

*k* = 第 *k* 个百分位数,

*n* = total number of data.

*n* = 数据的总数。

Expression for finding the percentile of a data value: $\left( \frac{x\text{~+~}0.5y}{n} \right)$(100)

求某一数据值百分位数的表达式:$\left( \frac{x\text{~+~}0.5y}{n} \right)$(100)

where *x* = the number of values counting from the bottom of the data list up to but not including the data value for which you want to find the percentile,

其中 *x* = 从数据列表底部向上数、直到(但不包含)你想求其百分位数的那个数据值之前的数据值个数,

*y* = the number of data values equal to the data value for which you want to find the percentile,

*y* = 与你想求其百分位数的那个数据值相等的数据值的个数,

*n* = total number of data

*n* = 数据的总数

2.5 Measures of the Center of the Data 2.5 数据中心位置的度量

$\mu = \frac{\sum{fm}}{\sum f}$ Where *f* = interval frequencies and *m* = interval midpoints.

$\mu = \frac{\sum{fm}}{\sum f}$,其中 *f* = 区间频数,*m* = 区间中点。

2.7 Measures of the Spread of the Data 2.7 数据离散程度的度量

$s_{x} = \sqrt{\frac{\sum{fm^{2}}}{n} - {\overline{x}}^{2}}$ where $\begin{array}{l}{s_{x} = \text{~sample~standard~deviation}} \\{\overline{x}\text{~=~sample~mean}}\end{array}$

$s_{x} = \sqrt{\frac{\sum{fm^{2}}}{n} - {\overline{x}}^{2}}$ 其中 $\begin{array}{l}{s_{x} = \text{~sample~standard~deviation}} \\{\overline{x}\text{~=~sample~mean}}\end{array}$

Practice 练习

2.1 Stem-and-Leaf Graphs (Stemplots), Line Graphs, and Bar Graphs 2.1 茎叶图(Stemplot)、线图与条形图

*For each of the following data sets, create a stem plot and identify any outliers.*

*对下列每个数据集,制作茎叶图并指出任何离群值。*

1. The miles per gallon rating for 30 cars are shown below (lowest to highest).

1. 30 辆汽车的每加仑英里数评级如下所示(从低到高)。

19, 19, 19, 20, 21, 21, 25, 25, 25, 26, 26, 28, 29, 31, 31, 32, 32, 33, 34, 35, 36, 37, 37, 38, 38, 38, 38, 41, 43, 43

19, 19, 19, 20, 21, 21, 25, 25, 25, 26, 26, 28, 29, 31, 31, 32, 32, 33, 34, 35, 36, 37, 37, 38, 38, 38, 38, 41, 43, 43

2\. The height in feet of 25 trees is shown below (lowest to highest).

2. 25 棵树的高度(英尺)如下所示(从低到高)。

25, 27, 33, 34, 34, 34, 35, 37, 37, 38, 39, 39, 39, 40, 41, 45, 46, 47, 49, 50, 50, 53, 53, 54, 54

25, 27, 33, 34, 34, 34, 35, 37, 37, 38, 39, 39, 39, 40, 41, 45, 46, 47, 49, 50, 50, 53, 53, 54, 54

3. The data are the prices of different laptops at an electronics store. Round each value to the nearest ten.

3. 数据是一家电子商店中不同笔记本电脑的价格。将每个数值四舍五入到最近的十位数。

249, 249, 260, 265, 265, 280, 299, 299, 309, 319, 325, 326, 350, 350, 350, 365, 369, 389, 409, 459, 489, 559, 569, 570, 610

249, 249, 260, 265, 265, 280, 299, 299, 309, 319, 325, 326, 350, 350, 350, 365, 369, 389, 409, 459, 489, 559, 569, 570, 610

4\. The data are daily high temperatures in a town for one month.

4. 数据是一个城镇某月每日的最高气温。

61, 61, 62, 64, 66, 67, 67, 67, 68, 69, 70, 70, 70, 71, 71, 72, 74, 74, 74, 75, 75, 75, 76, 76, 77, 78, 78, 79, 79, 95

61, 61, 62, 64, 66, 67, 67, 67, 68, 69, 70, 70, 70, 71, 71, 72, 74, 74, 74, 75, 75, 75, 76, 76, 77, 78, 78, 79, 79, 95

*For the next three exercises, use the data to construct a line graph.*

*对接下来三道习题,利用这些数据制作线图。*

5. In a survey, 40 people were asked how many times they visited a store before making a major purchase. The results are shown in Table 2.37.

5. 在一份调查中,40 人被问到在做出一项重大购买前光顾商店的次数。结果如表 2.37 所示。

| Number of times in store | Frequency |

| 光顾商店的次数 | 频数 |

|--------------------------|-----------|

|----------------|------|

| 1 | 4 |

| 1 | 4 |

| 2 | 10 |

| 2 | 10 |

| 3 | 16 |

| 3 | 16 |

| 4 | 6 |

| 4 | 6 |

| 5 | 4 |

| 5 | 4 |

Table 2.37

表 2.37

6. In a survey, several people were asked how many years it has been since they purchased a mattress. The results are shown in Table 2.38.

6. 在一份调查中,一些人被问到距他们上次购买床垫已经过去了多少年。结果如表 2.38 所示。

| Years since last purchase | Frequency |

| 距上次购买以来的年数 | 频数 |

|---------------------------|-----------|

|---------------------------|------|

| 0 | 2 |

| 0 | 2 |

| 1 | 8 |

| 1 | 8 |

| 2 | 13 |

| 2 | 13 |

| 3 | 22 |

| 3 | 22 |

| 4 | 16 |

| 4 | 16 |

| 5 | 9 |

| 5 | 9 |

Table 2.38

表 2.38

7. Several children were asked how many TV shows they watch each day. The results of the survey are shown in Table 2.39.

7. 一些儿童被问到他们每天观看多少个电视节目。调查结果如表 2.39 所示。

| Number of TV Shows | Frequency |

| 观看的电视节目数量 | 频数 |

|--------------------|-----------|

|--------------------|------|

| 0 | 12 |

| 0 | 12 |

| 1 | 18 |

| 1 | 18 |

| 2 | 36 |

| 2 | 36 |

| 3 | 7 |

| 3 | 7 |

| 4 | 2 |

| 4 | 2 |

Table 2.39

表 2.39

8. The students in Ms. Ramirez’s math class have birthdays in each of the four seasons. Table 2.40 shows the four seasons, the number of students who have birthdays in each season, and the percentage (%) of students in each group. Construct a bar graph showing the number of students.

8. Ramirez 女士数学班上的学生生日分布在四个季节。表 2.40 给出了四个季节、每个季节过生日的学生人数,以及每组学生所占的百分比(%)。制作一个显示学生人数的条形图。

| Seasons | Number of students | Proportion of population |

| 季节 | 学生数 | 占总体比例 |

|---------|--------------------|--------------------------|

|---------|--------------------|--------------------------|

| Spring | 8 | 24% |

| 春季 | 8 | 24% |

| Summer | 9 | 26% |

| 夏季 | 9 | 26% |

| Autumn | 11 | 32% |

| 秋季 | 11 | 32% |

| Winter | 6 | 18% |

| 冬季 | 6 | 18% |

Table 2.40

表 2.40

9. Using the data from Mrs. Ramirez’s math class supplied in Table 2.40, construct a bar graph showing the percentages.

9. 利用表 2.40 提供的 Ramirez 女士数学班的数据,制作一个显示百分比的条形图。

10\. David County has six high schools. Each school sent students to participate in a county-wide science competition. Table 2.41 shows the percentage breakdown of competitors from each school, and the percentage of the entire student population of the county that goes to each school. Construct a bar graph that shows the population percentage of competitors from each school.

10. David 县有六所高中。每所学校都派学生参加全县的科学竞赛。表 2.41 给出了每所学校参赛者的百分比构成,以及该县全体学生中就读于每所学校的百分比。制作一个条形图,显示每所学校参赛者在总体中所占的百分比。

| High School | Science competition population | Overall student population |

| 高中 | 科学竞赛人数占比 | 学生总体占比 |

|-------------|--------------------------------|----------------------------|

|-------------|--------------------------------|----------------------------|

| Alabaster | 28.9% | 8.6% |

| Alabaster | 28.9% | 8.6% |

| Concordia | 7.6% | 23.2% |

| Concordia | 7.6% | 23.2% |

| Genoa | 12.1% | 15.0% |

| Genoa | 12.1% | 15.0% |

| Mocksville | 18.5% | 14.3% |

| Mocksville | 18.5% | 14.3% |

| Tynneson | 24.2% | 10.1% |

| Tynneson | 24.2% | 10.1% |

| West End | 8.7% | 28.8% |

| West End | 8.7% | 28.8% |

Table 2.41

表 2.41

11. Use the data from the David County science competition supplied in Table 2.41. Construct a bar graph that shows the county-wide population percentage of students at each school.

11. 利用表 2.41 提供的 David 县科学竞赛数据。制作一个条形图,显示每所学校学生在全县总体中所占的百分比。

2.2 Histograms, Frequency Polygons, and Time Series Graphs 2.2 直方图、频数多边形与时间序列图

12\. Sixty-five randomly selected car salespersons were asked the number of cars they generally sell in one week. Fourteen people answered that they generally sell three cars; nineteen generally sell four cars; twelve generally sell five cars; nine generally sell six cars; eleven generally sell seven cars. Complete the table.

12. 65 名随机选取的汽车销售员被问到他们通常一周能卖出多少辆汽车。14 人回答通常卖 3 辆;19 人通常卖 4 辆;12 人通常卖 5 辆;9 人通常卖 6 辆;11 人通常卖 7 辆。补全表格。

| Data Value (# cars) | Frequency | Relative Frequency | Cumulative Relative Frequency |

| 数据值(汽车数量) | 频数 | 相对频数 | 累积相对频数 |

|---------------------|-----------|--------------------|-------------------------------|

|---------------------|-----------|--------------------|-------------------------------|

| | | | |

| | | | |

| | | | |

| | | | |

| | | | |

| | | | |

| | | | |

| | | | |

Table 2.42

表 2.42

13. What does the frequency column in Table 2.42 sum to? Why?

13. 表 2.42 中的频数一列求和等于多少?为什么?

14\. What does the relative frequency column in Table 2.42 sum to? Why?

14. 表 2.42 中的相对频数一列求和等于多少?为什么?

15. What is the difference between relative frequency and frequency for each data value in Table 2.42?

15. 表 2.42 中,每个数据值的相对频数与频数之间的区别是什么?

16\. What is the difference between cumulative relative frequency and relative frequency for each data value?

16. 对每个数据值而言,累积相对频数与相对频数之间的区别是什么?

17. To construct the histogram for the data in Table 2.42, determine appropriate minimum and maximum *x* and *y* values and the scaling. Sketch the histogram. Label the horizontal and vertical axes with words. Include numerical scaling.

17. 要为表 2.42 中的数据制作直方图,先确定合适的 *x* 和 *y* 的最小值与最大值以及标度。画出直方图的草图。用文字标注横轴和纵轴。包含数值标度。

18\. Construct a frequency polygon for the following:

18. 为下列数据制作频数多边形:

1. | Pulse Rates for Women | Frequency |

1. | 女性脉搏率 | 频数 |

|-----------------------|-----------|

|-----------------------|-----------|

| 60–69 | 12 |

| 60–69 | 12 |

| 70–79 | 14 |

| 70–79 | 14 |

| 80–89 | 11 |

| 80–89 | 11 |

| 90–99 | 1 |

| 90–99 | 1 |

| 100–109 | 1 |

| 100–109 | 1 |

| 110–119 | 0 |

| 110–119 | 0 |

| 120–129 | 1 |

| 120–129 | 1 |

Table 2.43

表 2.43

2. | Actual Speed in a 30 MPH Zone | Frequency |

2. | 30 英里/小时限速区内的实际车速 | 频数 |

|-------------------------------|-----------|

|-------------------------------|-----------|

| 42–45 | 25 |

| 42–45 | 25 |

| 46–49 | 14 |

| 46–49 | 14 |

| 50–53 | 7 |

| 50–53 | 7 |

| 54–57 | 3 |

| 54–57 | 3 |

| 58–61 | 1 |

| 58–61 | 1 |

Table 2.44

表 2.44

3. | Tar (mg) in Nonfiltered Cigarettes | Frequency |

3. | 未过滤香烟中的焦油(毫克) | 频数 |

|------------------------------------|-----------|

|------------------------------------|-----------|

| 10–13 | 1 |

| 10–13 | 1 |

| 14–17 | 0 |

| 14–17 | 0 |

| 18–21 | 15 |

| 18–21 | 15 |

| 22–25 | 7 |

| 22–25 | 7 |

| 26–29 | 2 |

| 26–29 | 2 |

Table 2.45

表 2.45

19. Construct a frequency polygon from the frequency distribution for the 50 highest ranked countries for depth of hunger.

19. 根据饥饿深度排名前 50 位国家的频数分布制作频数多边形。

| Depth of Hunger | Frequency |

| 饥饿深度 | 频数 |

|-----------------|-----------|

|-----------------|-----------|

| 230–259 | 21 |

| 230–259 | 21 |

| 260–289 | 13 |

| 260–289 | 13 |

| 290–319 | 5 |

| 290–319 | 5 |

| 320–349 | 7 |

| 320–349 | 7 |

| 350–379 | 1 |

| 350–379 | 1 |

| 380–409 | 1 |

| 380–409 | 1 |

| 410–439 | 1 |

| 410–439 | 1 |

Table 2.46

表 2.46

20. Use the two frequency tables to compare the life expectancy of men and women from 20 randomly selected countries. Include an overlayed frequency polygon and discuss the shapes of the distributions, the center, the spread, and any outliers. What can we conclude about the life expectancy of women compared to men?

20. 利用这两张频数表比较 20 个随机选取国家的男性和女性的预期寿命。包含一张叠加的频数多边形,并讨论分布的形状、中心、离散程度以及任何离群值。关于女性与男性的预期寿命,我们能得出什么结论?

| Life Expectancy at Birth – Women | Frequency |

| 出生时的预期寿命——女性 | 频数 |

|----------------------------------|-----------|

|----------------------------------|-----------|

| 49–55 | 3 |

| 49–55 | 3 |

| 56–62 | 3 |

| 56–62 | 3 |

| 63–69 | 1 |

| 63–69 | 1 |

| 70–76 | 3 |

| 70–76 | 3 |

| 77–83 | 8 |

| 77–83 | 8 |

| 84–90 | 2 |

| 84–90 | 2 |

Table 2.47

表 2.47

| Life Expectancy at Birth – Men | Frequency |

| 出生时的预期寿命——男性 | 频数 |

|--------------------------------|-----------|

|--------------------------------|-----------|

| 49–55 | 3 |

| 49–55 | 3 |

| 56–62 | 3 |

| 56–62 | 3 |

| 63–69 | 1 |

| 63–69 | 1 |

| 70–76 | 1 |

| 70–76 | 1 |

| 77–83 | 7 |

| 77–83 | 7 |

| 84–90 | 5 |

| 84–90 | 5 |

Table 2.48

表 2.48

21. Construct a times series graph for (a) the number of male births, (b) the number of female births, and (c) the total number of births.

21. 制作时间序列图,分别表示 (a) 男性出生数,(b) 女性出生数,(c) 出生总数。

| | | | | | | | |

| | | | | | | | |

|----------|--------|---------|---------|---------|---------|---------|---------|

|----------|--------|---------|---------|---------|---------|---------|---------|

| Sex/Year | 1855 | 1856 | 1857 | 1858 | 1859 | 1860 | 1861 |

| 性别/年份 | 1855 | 1856 | 1857 | 1858 | 1859 | 1860 | 1861 |

| Female | 45,545 | 49,582 | 50,257 | 50,324 | 51,915 | 51,220 | 52,403 |

| 女性 | 45,545 | 49,582 | 50,257 | 50,324 | 51,915 | 51,220 | 52,403 |

| Male | 47,804 | 52,239 | 53,158 | 53,694 | 54,628 | 54,409 | 54,606 |

| 男性 | 47,804 | 52,239 | 53,158 | 53,694 | 54,628 | 54,409 | 54,606 |

| Total | 93,349 | 101,821 | 103,415 | 104,018 | 106,543 | 105,629 | 107,009 |

| 总计 | 93,349 | 101,821 | 103,415 | 104,018 | 106,543 | 105,629 | 107,009 |

Table 2.49

表 2.49

| | | | | | | | | |

| | | | | | | | | |

|----------|---------|---------|---------|---------|---------|---------|---------|---------|

|----------|---------|---------|---------|---------|---------|---------|---------|---------|

| Sex/Year | 1862 | 1863 | 1864 | 1865 | 1866 | 1867 | 1868 | 1869 |

| 性别/年份 | 1862 | 1863 | 1864 | 1865 | 1866 | 1867 | 1868 | 1869 |

| Female | 51,812 | 53,115 | 54,959 | 54,850 | 55,307 | 55,527 | 56,292 | 55,033 |

| 女性 | 51,812 | 53,115 | 54,959 | 54,850 | 55,307 | 55,527 | 56,292 | 55,033 |

| Male | 55,257 | 56,226 | 57,374 | 58,220 | 58,360 | 58,517 | 59,222 | 58,321 |

| 男性 | 55,257 | 56,226 | 57,374 | 58,220 | 58,360 | 58,517 | 59,222 | 58,321 |

| Total | 107,069 | 109,341 | 112,333 | 113,070 | 113,667 | 114,044 | 115,514 | 113,354 |

| 总计 | 107,069 | 109,341 | 112,333 | 113,070 | 113,667 | 114,044 | 115,514 | 113,354 |

Table 2.50

表 2.50

| | | | | | | |

| | | | | | | |

|----------|---------|---------|---------|---------|---------|---------|

|----------|---------|---------|---------|---------|---------|---------|

| Sex/Year | 1870 | 1871 | 1872 | 1873 | 1874 | 1875 |

| 性别/年份 | 1870 | 1871 | 1872 | 1873 | 1874 | 1875 |

| Female | 56,431 | 56,099 | 57,472 | 58,233 | 60,109 | 60,146 |

| 女性 | 56,431 | 56,099 | 57,472 | 58,233 | 60,109 | 60,146 |

| Male | 58,959 | 60,029 | 61,293 | 61,467 | 63,602 | 63,432 |

| 男性 | 58,959 | 60,029 | 61,293 | 61,467 | 63,602 | 63,432 |

| Total | 115,390 | 116,128 | 118,765 | 119,700 | 123,711 | 123,578 |

| 总计 | 115,390 | 116,128 | 118,765 | 119,700 | 123,711 | 123,578 |

Table 2.51

表 2.51

22. The following data sets list full time police per 100,000 citizens along with homicides per 100,000 citizens for the city of Detroit, Michigan during the period from 1961 to 1973.

22. 以下数据集列出了密歇根州底特律市在 1961 年至 1973 年期间每 10 万市民中全职警察人数以及每 10 万市民中的凶杀案数。

| | | | | | | | |

| | | | | | | | |

|-----------|--------|-------|--------|--------|--------|--------|--------|

|-----------|--------|-------|--------|--------|--------|--------|--------|

| Year | 1961 | 1962 | 1963 | 1964 | 1965 | 1966 | 1967 |

| 年份 | 1961 | 1962 | 1963 | 1964 | 1965 | 1966 | 1967 |

| Police | 260.35 | 269.8 | 272.04 | 272.96 | 272.51 | 261.34 | 268.89 |

| 警察 | 260.35 | 269.8 | 272.04 | 272.96 | 272.51 | 261.34 | 268.89 |

| Homicides | 8.6 | 8.9 | 8.52 | 8.89 | 13.07 | 14.57 | 21.36 |

| 凶杀案 | 8.6 | 8.9 | 8.52 | 8.89 | 13.07 | 14.57 | 21.36 |

Table 2.52

表 2.52

| | | | | | | |

| | | | | | | |

|-----------|--------|--------|--------|--------|--------|--------|

|-----------|--------|--------|--------|--------|--------|--------|

| Year | 1968 | 1969 | 1970 | 1971 | 1972 | 1973 |

| 年份 | 1968 | 1969 | 1970 | 1971 | 1972 | 1973 |

| Police | 295.99 | 319.87 | 341.43 | 356.59 | 376.69 | 390.19 |

| 警察 | 295.99 | 319.87 | 341.43 | 356.59 | 376.69 | 390.19 |

| Homicides | 28.03 | 31.49 | 37.39 | 46.26 | 47.24 | 52.33 |

| 凶杀案 | 28.03 | 31.49 | 37.39 | 46.26 | 47.24 | 52.33 |

Table 2.53

表 2.53

1. Construct a double time series graph using a common *x*-axis for both sets of data.

1. 使用一条公共 *x* 轴为两组数据制作双时间序列图。

2. Which variable increased the fastest? Explain.

2. 哪个变量增长最快?请解释。

3. Did Detroit’s increase in police officers have an impact on the murder rate? Explain.

3. 底特律警察数量的增加是否对凶杀率产生了影响?请解释。

2.3 Measures of the Location of the Data 2.3 数据的位置度量

23. Listed are 29 ages for Academy Award winning best actors *in order from smallest to largest.*

23. 下面按从小到大顺序排列了 29 位奥斯卡最佳男主角获奖者的年龄。

18; 21; 22; 25; 26; 27; 29; 30; 31; 33; 36; 37; 41; 42; 47; 52; 55; 57; 58; 62; 64; 67; 69; 71; 72; 73; 74; 76; 77

18;21;22;25;26;27;29;30;31;33;36;37;41;42;47;52;55;57;58;62;64;67;69;71;72;73;74;76;77

1. Find the 40th percentile.

1. 求第 40 百分位数。

2. Find the 78th percentile.

2. 求第 78 百分位数。

24\. Listed are 32 ages for Academy Award winning best actors *in order from smallest to largest.*

24. 下面按从小到大顺序排列了 32 位奥斯卡最佳男主角获奖者的年龄。

18; 18; 21; 22; 25; 26; 27; 29; 30; 31; 31; 33; 36; 37; 37; 41; 42; 47; 52; 55; 57; 58; 62; 64; 67; 69; 71; 72; 73; 74; 76; 77

18;18;21;22;25;26;27;29;30;31;31;33;36;37;37;41;42;47;52;55;57;58;62;64;67;69;71;72;73;74;76;77

1. Find the percentile of 37.

1. 求 37 的百分位数。

2. Find the percentile of 72.

2. 求 72 的百分位数。

25. Jesse was ranked 37th in his graduating class of 180 students. At what percentile is Jesse’s ranking?

25. Jesse 在 180 名学生的毕业班中排名第 37。Jesse 的排名处于第几百分位数?

26\. 1. For runners in a race, a low time means a faster run. The winners in a race have the shortest running times. Is it more desirable to have a finish time with a high or a low percentile when running a race?

26. 1. 对赛跑者而言,时间越短跑得越快。赛跑中的获胜者用时最短。赛跑时,完成时间处于较高还是较低百分位更可取?

2. The 20th percentile of run times in a particular race is 5.2 minutes. Write a sentence interpreting the 20th percentile in the context of the situation.

2. 某场特定赛跑中跑步时间的第 20 百分位数是 5.2 分钟。写一句话,结合情境解释第 20 百分位数。

3. A bicyclist in the 90th percentile of a bicycle race completed the race in 1 hour and 12 minutes. Is he among the fastest or slowest cyclists in the race? Write a sentence interpreting the 90th percentile in the context of the situation.

3. 一名在自行车赛中位列第 90 百分位数的骑行者用时 1 小时 12 分钟完成比赛。他是该比赛中最快还是最慢的骑行者之一?写一句话,结合情境解释第 90 百分位数。

27. 1. For runners in a race, a higher speed means a faster run. Is it more desirable to have a speed with a high or a low percentile when running a race?

27. 1. 对赛跑者而言,速度越快跑得越快。赛跑时,速度处于较高还是较低百分位更可取?

2. The 40th percentile of speeds in a particular race is 7.5 miles per hour. Write a sentence interpreting the 40th percentile in the context of the situation.

2. 某场特定赛跑中速度的第 40 百分位数是每小时 7.5 英里。写一句话,结合情境解释第 40 百分位数。

28\. On an exam, would it be more desirable to earn a grade with a high or low percentile? Explain.

28. 在考试中,获得较高百分位还是较低百分位的成绩更可取?请解释。

29. Mina is waiting in line at the Department of Motor Vehicles (DMV). Her wait time of 32 minutes is the 85th percentile of wait times. Is that good or bad? Write a sentence interpreting the 85th percentile in the context of this situation.

29. Mina 在机动车辆管理部(DMV)排队等候。她 32 分钟的等候时间处于等候时间的第 85 百分位数。这是好事还是坏事?写一句话,结合此情境解释第 85 百分位数。

30\. In a survey collecting data about the salaries earned by recent college graduates, Li found that her salary was in the 78th percentile. Should Li be pleased or upset by this result? Explain.

30. 在一项收集近期大学毕业生薪资数据的调查中,Li 发现自己的薪资处于第 78 百分位数。Li 应该对此结果感到高兴还是沮丧?请解释。

31. In a study collecting data about the repair costs of damage to automobiles in a certain type of crash tests, a certain model of car had \$1,700 in damage and was in the 90th percentile. Should the manufacturer and the consumer be pleased or upset by this result? Explain and write a sentence that interprets the 90th percentile in the context of this problem.

31. 在一项收集某类碰撞测试中汽车损坏维修费用数据的研究中,某款汽车的损坏费用为 1700 美元,处于第 90 百分位数。制造商和消费者应该对此结果感到高兴还是沮丧?请解释,并写一句话,结合此问题情境解释第 90 百分位数。

32\. The University of California has two criteria used to set admission standards for freshman to be admitted to a college in the UC system:

32. 加利福尼亚大学采用两项标准来设定 UC 系统各学院新生的录取标准:

1. Students' GPAs and scores on standardized tests (SATs and ACTs) are entered into a formula that calculates an "admissions index" score. The admissions index score is used to set eligibility standards intended to meet the goal of admitting the top 12% of high school students in the state. In this context, what percentile does the top 12% represent?

1. 学生的 GPA 和标准化考试(SAT 和 ACT)成绩被代入一个公式,计算出“录取指数”分数。该录取指数分数用于设定资格标准,以达到录取本州前 12% 高中生的目标。在此情境下,前 12% 代表第几百分位数?

2. Students whose GPAs are at or above the 96th percentile of all students at their high school are eligible (called eligible in the local context), even if they are not in the top 12% of all students in the state. What percentage of students from each high school are "eligible in the local context"?

2. 在本校所有学生中 GPA 处于或高于第 96 百分位数的学生具备资格(称为“本地语境下的合格”),即使他们不在本州所有学生的前 12% 之列。每所高中中,百分之多少的学生属于“本地语境下的合格”?

33. Suppose that you are buying a house. You and your realtor have determined that the most expensive house you can afford is the 34th percentile. The 34th percentile of housing prices is \$240,000 in the town you want to move to. In this town, can you afford 34% of the houses or 66% of the houses?

33. 假设你要买房。你和房产经纪人已确定,你买得起的最贵房子是第 34 百分位数。在你想要搬去的城镇,房价的第 34 百分位数是 240,000 美元。在这个城镇,你买得起 34% 还是 66% 的房子?

Use the following information to answer the next six exercises. Sixty-five randomly selected car salespersons were asked the number of cars they generally sell in one week. Fourteen people answered that they generally sell three cars; nineteen generally sell four cars; twelve generally sell five cars; nine generally sell six cars; eleven generally sell seven cars.

利用以下信息回答接下来六道习题。65 名随机选取的汽车销售员被问到他们通常一周能卖出多少辆汽车。14 人回答通常卖 3 辆;19 人通常卖 4 辆;12 人通常卖 5 辆;9 人通常卖 6 辆;11 人通常卖 7 辆。

34\. First quartile = \_\_\_\_\_\_\_

34. 第一四分位数 = \_\_\_\_\_\_\_

35. Second quartile = median = 50th percentile = \_\_\_\_\_\_\_

35. 第二四分位数 = 中位数 = 第 50 百分位数 = \_\_\_\_\_\_\_

36\. Third quartile = \_\_\_\_\_\_\_

36. 第三四分位数 = \_\_\_\_\_\_\_

37. Interquartile range (*IQR*) = \_\_\_\_\_ – \_\_\_\_\_ = \_\_\_\_\_

37. 四分位距(*IQR*)= \_\_\_\_\_ – \_\_\_\_\_ = \_\_\_\_\_

38\. 10th percentile = \_\_\_\_\_\_\_

38. 第 10 百分位数 = \_\_\_\_\_\_\_

39. 70th percentile = \_\_\_\_\_\_\_

39. 第 70 百分位数 = \_\_\_\_\_\_\_

2.4 Box Plots 2.4 箱线图

Use the following information to answer the next two exercises. Sixty-five randomly selected car salespersons were asked the number of cars they generally sell in one week. Fourteen people answered that they generally sell three cars; nineteen generally sell four cars; twelve generally sell five cars; nine generally sell six cars; eleven generally sell seven cars.

用下列信息回答接下来两道习题。随机选取的 65 名汽车销售人员被问及他们通常一周能售出多少辆汽车。14 人说通常售出 3 辆;19 人通常售出 4 辆;12 人通常售出 5 辆;9 人通常售出 6 辆;11 人通常售出 7 辆。

40\. Construct a box plot below. Use a ruler to measure and scale accurately.

40. 在下方绘制一个箱线图。使用直尺准确测量并按比例作图。

41. Looking at your box plot, does it appear that the data are concentrated together, spread out evenly, or concentrated in some areas, but not in others? How can you tell?

41. 观察你的箱线图,数据看起来是集中在一起、均匀分布,还是集中在部分区域而非全部区域?你如何判断?

2.5 Measures of the Center of the Data 2.5 数据的中心度量

42\. Find the mean for the following frequency tables.

42. 求下列频数表的均值。

1. | Grade | Frequency |

1. | 成绩 | 频数 |

|-----------|-----------|

|-----------|-----------|

| 49.5–59.5 | 2 |

| 49.5–59.5 | 2 |

| 59.5–69.5 | 3 |

| 59.5–69.5 | 3 |

| 69.5–79.5 | 8 |

| 69.5–79.5 | 8 |

| 79.5–89.5 | 12 |

| 79.5–89.5 | 12 |

| 89.5–99.5 | 5 |

| 89.5–99.5 | 5 |

Table 2.54

表 2.54

2. | Daily Low Temperature | Frequency |

2. | 每日最低气温 | 频数 |

|-----------------------|-----------|

|-----------------------|-----------|

| 49.5–59.5 | 53 |

| 49.5–59.5 | 53 |

| 59.5–69.5 | 32 |

| 59.5–69.5 | 32 |

| 69.5–79.5 | 15 |

| 69.5–79.5 | 15 |

| 79.5–89.5 | 1 |

| 79.5–89.5 | 1 |

| 89.5–99.5 | 0 |

| 89.5–99.5 | 0 |

Table 2.55

表 2.55

3. | Points per Game | Frequency |

3. | 每场得分 | 频数 |

|-----------------|-----------|

|-----------------|-----------|

| 49.5–59.5 | 14 |

| 49.5–59.5 | 14 |

| 59.5–69.5 | 32 |

| 59.5–69.5 | 32 |

| 69.5–79.5 | 15 |

| 69.5–79.5 | 15 |

| 79.5–89.5 | 23 |

| 79.5–89.5 | 23 |

| 89.5–99.5 | 2 |

| 89.5–99.5 | 2 |

Table 2.56

表 2.56

*Use the following information to answer the next three exercises:* The following data show the lengths of boats moored in a marina. The data are ordered from smallest to largest: 16; 17; 19; 20; 20; 21; 23; 24; 25; 25; 25; 26; 26; 27; 27; 27; 28; 29; 30; 32; 33; 33; 34; 35; 37; 39; 40

*用下列信息回答接下来三道习题:* 下列数据显示了停泊在船坞中的船只长度。数据按从小到大排列:16;17;19;20;20;21;23;24;25;25;25;26;26;27;27;27;28;29;30;32;33;33;34;35;37;39;40

43. Calculate the mean.

43. 计算均值。

44\. Identify the median.

44. 确定中位数。

45. Identify the mode.

45. 确定众数。

*Use the following information to answer the next three exercises:* Sixty-five randomly selected car salespersons were asked the number of cars they generally sell in one week. Fourteen people answered that they generally sell three cars; nineteen generally sell four cars; twelve generally sell five cars; nine generally sell six cars; eleven generally sell seven cars. Calculate the following:

*用下列信息回答接下来三道习题:* 随机选取的 65 名汽车销售人员被问及他们通常一周能售出多少辆汽车。14 人说通常售出 3 辆;19 人通常售出 4 辆;12 人通常售出 5 辆;9 人通常售出 6 辆;11 人通常售出 7 辆。计算下列各项:

46\. sample mean = $\overline{x}$ = \_\_\_\_\_\_\_

46. 样本均值 = $\overline{x}$ = \_\_\_\_\_\_\_

47. median = \_\_\_\_\_\_\_

47. 中位数 = \_\_\_\_\_\_\_

48\. mode = \_\_\_\_\_\_\_

48. 众数 = \_\_\_\_\_\_\_

2.6 Skewness and the Mean, Median, and Mode 2.6 偏度与均值、中位数和众数

*Use the following information to answer the next three exercises:* State whether the data are symmetrical, skewed to the left, or skewed to the right.

*用下列信息回答接下来三道习题:* 指出数据是对称的、左偏还是右偏。

49. 1; 1; 1; 2; 2; 2; 2; 3; 3; 3; 3; 3; 3; 3; 3; 4; 4; 4; 5; 5

49. 1;1;1;2;2;2;2;3;3;3;3;3;3;3;3;4;4;4;5;5

50\. 16; 17; 19; 22; 22; 22; 22; 22; 23

50. 16;17;19;22;22;22;22;22;23

51. 87; 87; 87; 87; 87; 88; 89; 89; 90; 91

51. 87;87;87;87;87;88;89;89;90;91

52\. When the data are skewed left, what is the typical relationship between the mean and median?

52. 当数据左偏时,均值与中位数之间典型的关系是什么?

53. When the data are symmetrical, what is the typical relationship between the mean and median?

53. 当数据对称时,均值与中位数之间典型的关系是什么?

54\. What word describes a distribution that has two modes?

54. 哪个词用来描述具有两个众数的分布?

55. Describe the shape of this distribution.

55. 描述该分布的形状。

56\. Describe the relationship between the mode and the median of this distribution.

56. 描述该分布的众数与中位数之间的关系。

57. Describe the relationship between the mean and the median of this distribution.

57. 描述该分布的均值与中位数之间的关系。

58\. Describe the shape of this distribution.

58. 描述该分布的形状。

59. Describe the relationship between the mode and the median of this distribution.

59. 描述该分布的众数与中位数之间的关系。

60\. Are the mean and the median the exact same in this distribution? Why or why not?

60. 在该分布中,均值与中位数是否完全相同?为什么相同或为什么不同?

61. Describe the shape of this distribution.

61. 描述该分布的形状。

62\. Describe the relationship between the mode and the median of this distribution.

62. 描述该分布的众数与中位数之间的关系。

63. Describe the relationship between the mean and the median of this distribution.

63. 描述该分布的均值与中位数之间的关系。

64\. The mean and median for the data are the same. 3; 4; 5; 5; 6; 6; 6; 6; 7; 7; 7; 7; 7; 7; 7 Is the data perfectly symmetrical? Why or why not?

64. 该数据的均值与中位数相同。3;4;5;5;6;6;6;6;7;7;7;7;7;7;7 该数据是否完全对称?为什么是或为什么不是?

65. Which is the greatest, the mean, the mode, or the median of the data set? 11; 11; 12; 12; 12; 12; 13; 15; 17; 22; 22; 22

65. 在该数据集中,均值、众数还是中位数最大?11;11;12;12;12;12;13;15;17;22;22;22

66\. Which is the least, the mean, the mode, and the median of the data set? 56; 56; 56; 58; 59; 60; 62; 64; 64; 65; 67

66. 在该数据集中,均值、众数和中位数中哪个最小?56;56;56;58;59;60;62;64;64;65;67

67. Of the three measures, which tends to reflect skewing the most, the mean, the mode, or the median? Why?

67. 在这三个度量中,哪个最容易反映偏斜——均值、众数还是中位数?为什么?

68\. In a perfectly symmetrical distribution, when would the mode be different from the mean and median?

68. 在一个完全对称的分布中,众数何时会不同于均值和中位数?

2.7 Measures of the Spread of the Data 2.7 数据的离散程度度量

*Use the following information to answer the next two exercises*: The following data are the distances between 20 retail stores and a large distribution center. The distances are in miles.

*用下列信息回答接下来两道习题*:下列数据是 20 家零售店与一个大配送中心之间的距离。距离以英里计。

29; 37; 38; 40; 58; 67; 68; 69; 76; 86; 87; 95; 96; 96; 99; 106; 112; 127; 145; 150

29;37;38;40;58;67;68;69;76;86;87;95;96;96;99;106;112;127;145;150

69. Use a graphing calculator or computer to find the standard deviation and round to the nearest tenth.

69. 使用图形计算器或计算机求标准差,并四舍五入到十分位。

70\. Find the value that is one standard deviation below the mean.

70. 求低于均值一个标准差的值。

71. Two baseball players, Fredo and Karl, on different teams wanted to find out who had the higher batting average when compared to his team. Which baseball player had the higher batting average when compared to his team?

71. 两名分属不同球队的棒球运动员 Fredo 和 Karl 想弄清楚,与各自球队相比,谁的击球率更高。哪位棒球运动员与各自球队相比击球率更高?

| Baseball Player | Batting Average | Team Batting Average | Team Standard Deviation |

| 棒球运动员 | 击球率 | 球队击球率 | 球队标准差 |

|-----------------|-----------------|----------------------|-------------------------|

|-----------------|-----------------|----------------------|-------------------------|

| Fredo | 0.158 | 0.166 | 0.012 |

| Fredo | 0.158 | 0.166 | 0.012 |

| Karl | 0.177 | 0.189 | 0.015 |

| Karl | 0.177 | 0.189 | 0.015 |

Table 2.57

表 2.57

72\. Use Table 2.57 to find the value that is three standard deviations:

72. 使用表 2.57 求三个标准差的值:

*Find the standard deviation for the following frequency tables using the formula. Check the calculations with the TI 83/84*.

*用下列公式计算下列频数表的标准差。用 TI 83/84 检验计算结果。*

73\. Find the standard deviation for the following frequency tables using the formula. Check the calculations with the TI 83/84.

73. 用下列公式计算下列频数表的标准差。用 TI 83/84 检验计算结果。

1. | Grade | Frequency |

1. | 成绩 | 频数 |

|-----------|-----------|

|-----------|-----------|

| 49.5–59.5 | 2 |

| 49.5–59.5 | 2 |

| 59.5–69.5 | 3 |

| 59.5–69.5 | 3 |

| 69.5–79.5 | 8 |

| 69.5–79.5 | 8 |

| 79.5–89.5 | 12 |

| 79.5–89.5 | 12 |

| 89.5–99.5 | 5 |

| 89.5–99.5 | 5 |

Table 2.58

表 2.58

2. | Daily Low Temperature | Frequency |

2. | 每日最低气温 | 频数 |

|-----------------------|-----------|

|-----------------------|-----------|

| 49.5–59.5 | 53 |

| 49.5–59.5 | 53 |

| 59.5–69.5 | 32 |

| 59.5–69.5 | 32 |

| 69.5–79.5 | 15 |

| 69.5–79.5 | 15 |

| 79.5–89.5 | 1 |

| 79.5–89.5 | 1 |

| 89.5–99.5 | 0 |

| 89.5–99.5 | 0 |

Table 2.59

表 2.59

3. | Points per Game | Frequency |

3. | 每场得分 | 频数 |

|-----------------|-----------|

|-----------------|-----------|

| 49.5–59.5 | 14 |

| 49.5–59.5 | 14 |

| 59.5–69.5 | 32 |

| 59.5–69.5 | 32 |

| 69.5–79.5 | 15 |

| 69.5–79.5 | 15 |

| 79.5–89.5 | 23 |

| 79.5–89.5 | 23 |

| 89.5–99.5 | 2 |

| 89.5–99.5 | 2 |

Table 2.60

表 2.60

Homework 作业

2.1 Stem-and-Leaf Graphs (Stemplots), Line Graphs, and Bar Graphs 2.1 茎叶图(茎叶图)、线图与条形图

74\. Student grades on a chemistry exam were: 77, 78, 76, 81, 86, 51, 79, 82, 84, 99

74. 一次化学考试的学生成绩为:77、78、76、81、86、51、79、82、84、99

1. Construct a stem-and-leaf plot of the data.

1. 为这些数据构造一个茎叶图。

2. Are there any potential outliers? If so, which scores are they? Why do you consider them outliers?

2. 是否存在潜在的离群值?如果有,是哪些分数?你为什么认为它们是离群值?

75. Table 2.61 contains the 2010 obesity rates in U.S. states and Washington, DC.

75. 表 2.61 包含了 2010 年美国各州及华盛顿特区的肥胖率。

| State | Percent (%) | State | Percent (%) | State | Percent (%) |

| 州 | 百分比(%) | 州 | 百分比(%) | 州 | 百分比(%) |

|----------------|-------------|----------------|-------------|----------------|-------------|

|----------------|-------------|----------------|-------------|----------------|-------------|

| Alabama | 32.2 | Kentucky | 31.3 | North Dakota | 27.2 |

| Alabama | 32.2 | Kentucky | 31.3 | North Dakota | 27.2 |

| Alaska | 24.5 | Louisiana | 31.0 | Ohio | 29.2 |

| Alaska | 24.5 | Louisiana | 31.0 | Ohio | 29.2 |

| Arizona | 24.3 | Maine | 26.8 | Oklahoma | 30.4 |

| Arizona | 24.3 | Maine | 26.8 | Oklahoma | 30.4 |

| Arkansas | 30.1 | Maryland | 27.1 | Oregon | 26.8 |

| Arkansas | 30.1 | Maryland | 27.1 | Oregon | 26.8 |

| California | 24.0 | Massachusetts | 23.0 | Pennsylvania | 28.6 |

| California | 24.0 | Massachusetts | 23.0 | Pennsylvania | 28.6 |

| Colorado | 21.0 | Michigan | 30.9 | Rhode Island | 25.5 |

| Colorado | 21.0 | Michigan | 30.9 | Rhode Island | 25.5 |

| Connecticut | 22.5 | Minnesota | 24.8 | South Carolina | 31.5 |

| Connecticut | 22.5 | Minnesota | 24.8 | South Carolina | 31.5 |

| Delaware | 28.0 | Mississippi | 34.0 | South Dakota | 27.3 |

| Delaware | 28.0 | Mississippi | 34.0 | South Dakota | 27.3 |

| Washington, DC | 22.2 | Missouri | 30.5 | Tennessee | 30.8 |

| Washington, DC | 22.2 | Missouri | 30.5 | Tennessee | 30.8 |

| Florida | 26.6 | Montana | 23.0 | Texas | 31.0 |

| Florida | 26.6 | Montana | 23.0 | Texas | 31.0 |

| Georgia | 29.6 | Nebraska | 26.9 | Utah | 22.5 |

| Georgia | 29.6 | Nebraska | 26.9 | Utah | 22.5 |

| Hawaii | 22.7 | Nevada | 22.4 | Vermont | 23.2 |

| Hawaii | 22.7 | Nevada | 22.4 | Vermont | 23.2 |

| Idaho | 26.5 | New Hampshire | 25.0 | Virginia | 26.0 |

| Idaho | 26.5 | New Hampshire | 25.0 | Virginia | 26.0 |

| Illinois | 28.2 | New Jersey | 23.8 | Washington | 25.5 |

| Illinois | 28.2 | New Jersey | 23.8 | Washington | 25.5 |

| Indiana | 29.6 | New Mexico | 25.1 | West Virginia | 32.5 |

| Indiana | 29.6 | New Mexico | 25.1 | West Virginia | 32.5 |

| Iowa | 28.4 | New York | 23.9 | Wisconsin | 26.3 |

| Iowa | 28.4 | New York | 23.9 | Wisconsin | 26.3 |

| Kansas | 29.4 | North Carolina | 27.8 | Wyoming | 25.1 |

| Kansas | 29.4 | North Carolina | 27.8 | Wyoming | 25.1 |

Table 2.61

表 2.61

1. Use a random number generator to randomly pick eight states. Construct a bar graph of the obesity rates of those eight states.

1. 使用随机数生成器随机选取八个州。为这八个州的肥胖率绘制一个条形图。

2. Construct a bar graph for all the states beginning with the letter "A."

2. 为所有以字母 "A" 开头的州绘制一个条形图。

3. Construct a bar graph for all the states beginning with the letter "M."

3. 为所有以字母 "M" 开头的州绘制一个条形图。

2.2 Histograms, Frequency Polygons, and Time Series Graphs 2.2 直方图、频数多边形与时间序列图

76\.

76.

Suppose that three book publishers were interested in the number of fiction paperbacks adult consumers purchase per month. Each publisher conducted a survey. In the survey, adult consumers were asked the number of fiction paperbacks they had purchased the previous month. The results are as follows:

假设三家图书出版商对成年消费者每月购买的平装小说数量感兴趣。每家出版商都进行了一项调查。调查中,成年消费者被问及他们在上个月购买的平装小说数量。结果如下:
Table 2.62 Publisher A
\# of booksFreq.Rel. Freq.
010
112
216
312
48
56
62
82
表 2.62 出版商 A
购书数量频数相对频数
010
112
216
312
48
56
62
82
Table 2.63 Publisher B
\# of booksFreq.Rel. Freq.
018
124
224
322
415
510
75
91
表 2.63 出版商 B
购书数量频数相对频数
018
124
224
322
415
510
75
91
Table 2.64 Publisher C
\# of booksFreq.Rel. Freq.
0–120
2–335
4–512
6–72
8–91
表 2.64 出版商 C
购书数量频数相对频数
0–120
2–335
4–512
6–72
8–91

1. Find the relative frequencies for each survey. Write them in the charts.

1. 求出每项调查的相对频数,并填入表格中。

2. Using either a graphing calculator, computer, or by hand, use the frequency column to construct a histogram for each publisher's survey. For Publishers A and B, make bar widths of one. For Publisher C, make bar widths of two.

2. 使用图形计算器、计算机或手工,利用频数一列为每家出版商的调查绘制直方图。对出版商 A 和 B,令组宽为 1;对出版商 C,令组宽为 2。

3. In complete sentences, give two reasons why the graphs for Publishers A and B are not identical.

3. 用完整的句子,给出出版商 A 与 B 的图形不相同的两条理由。

4. Would you have expected the graph for Publisher C to look like the other two graphs? Why or why not?

4. 你是否预期出版商 C 的图形会与该两张图形相似?为什么(或为什么不)?

5. Make new histograms for Publisher A and Publisher B. This time, make bar widths of two.

5. 为出版商 A 和 B 重新绘制直方图。这次令组宽为 2。

6. Now, compare the graph for Publisher C to the new graphs for Publishers A and B. Are the graphs more similar or more different? Explain your answer.

6. 现在,将出版商 C 的图形与出版商 A、B 的新图形进行比较。这些图形是更相似还是更不同?解释你的答案。

77.

77.

Often, cruise ships conduct all on-board transactions, with the exception of gambling, on a cashless basis. At the end of the cruise, guests pay one bill that covers all onboard transactions. Suppose that 60 single travelers and 70 couples were surveyed as to their on-board bills for a seven-day cruise from Los Angeles to the Mexican Riviera. Following is a summary of the bills for each group.

邮轮通常(赌博除外)以无现金方式进行所有船上交易。邮轮结束时,乘客支付一张涵盖所有船上交易的账单。假设调查了 60 名单身旅行者和 70 对夫妇,询问他们在从洛杉矶到墨西哥里维耶拉的七日邮轮行程中的船上消费金额。以下是各组消费的汇总。
Table 2.65 Singles
Amount(\$)FrequencyRel. Frequency
51–1005
101–15010
151–20015
201–25015
251–30010
301–3505
表 2.65 单身者
金额(美元)频数相对频数
51–1005
101–15010
151–20015
201–25015
251–30010
301–3505
Table 2.66 Couples
Amount(\$)FrequencyRel. Frequency
100–1505
201–2505
251–3005
301–3505
351–40010
401–45010
451–50010
501–55010
551–6005
601–6505
表 2.66 夫妇
金额(美元)频数相对频数
100–1505
201–2505
251–3005
301–3505
351–40010
401–45010
451–50010
501–55010
551–6005
601–6505

1. Fill in the relative frequency for each group.

1. 填入各组的相对频数。

2. Construct a histogram for the singles group. Scale the *x*-axis by \$50 widths. Use relative frequency on the *y*-axis.

2. 为单身者组绘制直方图。*x* 轴以 50 美元为组宽。*y* 轴使用相对频数。

3. Construct a histogram for the couples group. Scale the *x*-axis by \$50 widths. Use relative frequency on the *y*-axis.

3. 为夫妇组绘制直方图。*x* 轴以 50 美元为组宽。*y* 轴使用相对频数。

4. Compare the two graphs:

4. 比较这两张图:

1. List two similarities between the graphs.

1. 列出两张图形的两个相似之处。

2. List two differences between the graphs.

2. 列出两张图形的两个不同之处。

3. Overall, are the graphs more similar or different?

3. 总体而言,两张图形是更相似还是更不同?

5. Construct a new graph for the couples by hand. Since each couple is paying for two individuals, instead of scaling the *x*-axis by \$50, scale it by \$100. Use relative frequency on the *y*-axis.

5. 手工为夫妇组重新绘制一张图。由于每对夫妇支付两个人的费用,*x* 轴不按 50 美元、而按 100 美元为组宽。*y* 轴使用相对频数。

6. Compare the graph for the singles with the new graph for the couples:

6. 将单身者组的图与夫妇组的新图进行比较:

1. List two similarities between the graphs.

1. 列出两张图形的两个相似之处。

2. Overall, are the graphs more similar or different?

2. 总体而言,两张图形是更相似还是更不同?

7. How did scaling the couples graph differently change the way you compared it to the singles graph?

7. 对夫妇组的图形采用不同的组宽刻度,如何改变了你将其与单身者组图形比较的方式?

8. Based on the graphs, do you think that individuals spend the same amount, more or less, as singles as they do person by person as a couple? Explain why in one or two complete sentences.

8. 根据图形,你认为个人作为单身者时的消费金额,与作为夫妇中一员时逐人消费的金额相比,是相同、更多还是更少?用一两句完整的话解释原因。

78\.

78.

Twenty-five randomly selected students were asked the number of movies they watched the previous week. The results are as follows.

随机选取的 25 名学生被问及他们在上一周观看的电影数量。结果如下:
Table 2.67
\# of moviesFrequencyRelative FrequencyCumulative Relative Frequency
05
19
26
34
41
表 2.67
电影数量频数相对频数累积相对频数
05
19
26
34
41

1. Construct a histogram of the data.

1. 根据数据绘制直方图。

2. Complete the columns of the chart.

2. 补全表格的各列。

*Use the following information to answer the next two exercises:* Suppose one hundred eleven people who shopped in a special t-shirt store were asked the number of t-shirts they own costing more than \$19 each.

(用以下信息回答接下来的两道习题:)假设在一家特卖 T 恤店里购物的 111 人被问及他们拥有的、单价超过 19 美元的 T 恤数量。

79.

79.

The percentage of people who own at most three t-shirts costing more than \$19 each is approximately:

拥有至多三件单价超过 19 美元的 T 恤的人,其百分比约为:

1. 21

1. 21

2. 59

2. 59

3. 41

3. 41

4. Cannot be determined

4. 无法确定

80\.

80.

If the data were collected by asking the first 111 people who entered the store, then the type of sampling is:

如果数据是通过询问进入商店的前 111 人收集的,那么抽样类型是:

1. cluster

1. 整群抽样

2. simple random

2. 简单随机抽样

3. stratified

3. 分层抽样

4. convenience

4. 方便抽样

81.

81.

Following are the 2010 obesity rates by U.S. states and Washington, DC.

以下是 2010 年美国各州及华盛顿特区的肥胖率。
Table 2.68
StatePercent (%)StatePercent (%)StatePercent (%)
Alabama32.2Kentucky31.3North Dakota27.2
Alaska24.5Louisiana31.0Ohio29.2
Arizona24.3Maine26.8Oklahoma30.4
Arkansas30.1Maryland27.1Oregon26.8
California24.0Massachusetts23.0Pennsylvania28.6
Colorado21.0Michigan30.9Rhode Island25.5
Connecticut22.5Mississippi34.0South Carolina31.5
Delaware28.0Missouri30.5South Dakota27.3
Washington, DC22.2Nebraska26.9Tennessee30.8
Florida26.6Montana23.0Texas31.0
Georgia29.6Nevada22.4Utah22.5
Hawaii22.7New Hampshire25.0Vermont23.2
Idaho26.5New Jersey23.8Virginia26.0
Illinois28.2New Mexico25.1Washington25.5
Indiana29.6New York23.9West Virginia32.5
Iowa28.4North Carolina27.8Wisconsin26.3
Kansas29.4North Carolina27.8Wyoming25.1
表 2.68
百分比(%)百分比(%)百分比(%)
Alabama32.2Kentucky31.3North Dakota27.2
Alaska24.5Louisiana31.0Ohio29.2
Arizona24.3Maine26.8Oklahoma30.4
Arkansas30.1Maryland27.1Oregon26.8
California24.0Massachusetts23.0Pennsylvania28.6
Colorado21.0Michigan30.9Rhode Island25.5
Connecticut22.5Mississippi34.0South Carolina31.5
Delaware28.0Missouri30.5South Dakota27.3
Washington, DC22.2Nebraska26.9Tennessee30.8
Florida26.6Montana23.0Texas31.0
Georgia29.6Nevada22.4Utah22.5
Hawaii22.7New Hampshire25.0Vermont23.2
Idaho26.5New Jersey23.8Virginia26.0
Illinois28.2New Mexico25.1Washington25.5
Indiana29.6New York23.9West Virginia32.5
Iowa28.4North Carolina27.8Wisconsin26.3
Kansas29.4North Carolina27.8Wyoming25.1

Construct a bar graph of obesity rates of your state and the four states closest to your state. Hint: Label the *x*-axis with the states.

为你所在州及离你所在州最近的四个州绘制肥胖率的条形图。提示:用各州名称标注 *x* 轴。

2.3 Measures of the Location of the Data 2.3 数据的位置度量

82\.

82.

The median age for Black people in the U.S. currently is 30.9 years; for U.S. White people it is 42.3 years.

目前美国黑人的年龄中位数为 30.9 岁;美国白人的年龄中位数为 42.3 岁。

1. Based upon this information, give two reasons why the Black median age could be lower than the White median age.

1. 基于这一信息,给出黑人年龄中位数可能低于白人年龄中位数的两条理由。

2. Does the lower median age for Black people necessarily mean that Blackpeople die younger than White people? Why or why not?

2. 黑人较低的年龄中位数是否必然意味着黑人比白人去世得更早?为什么(或为什么不)?

3. How might it be possible for Black people and White people to die at approximately the same age, but for the median age for White people to be higher?

3. 黑人与白人可能在几乎相同的年龄去世,但白人的年龄中位数却更高,这怎么可能?

83.

83.

Six hundred adult Americans were asked by telephone poll, "What do you think constitutes a middle-class income?" The results are in Table 2.69. Also, include left endpoint, but not the right endpoint.

通过电话民调询问 600 名美国成年人:"你认为怎样的收入算中产阶级?"结果见表 2.69。此外,区间包括左端点,但不包括右端点。
Table 2.69
Salary (\$)Relative Frequency
< 20,0000.02
20,000–25,0000.09
25,000–30,0000.19
30,000–40,0000.26
40,000–50,0000.18
50,000–75,0000.17
75,000–99,9990.02
100,000+0.01
表 2.69
工资(美元)相对频数
< 20,0000.02
20,000–25,0000.09
25,000–30,0000.19
30,000–40,0000.26
40,000–50,0000.18
50,000–75,0000.17
75,000–99,9990.02
100,000+0.01

1. What percentage of the survey answered "not sure"?

1. 调查中回答"不确定"的百分比是多少?

2. What percentage think that middle-class is from \$25,000 to \$50,000?

2. 认为中产阶级收入在 25,000 美元到 50,000 美元之间的人占百分之多少?

3. Construct a histogram of the data.

3. 根据数据绘制直方图。

1. Should all bars have the same width, based on the data? Why or why not?

1. 根据数据,所有条形是否应具有相同的宽度?为什么(或为什么不)?

2. How should the \<20,000 and the 100,000+ intervals be handled? Why?

2. 应如何处理 < 20,000 和 100,000+ 这两个区间?为什么?

4. Find the 40th and 80th percentiles

4. 求第 40th 和第 80th 百分位数

5. Construct a bar graph of the data

5. 根据数据绘制条形图

84\.

84.

Given the following box plot:

给定如下箱线图:

1. which quarter has the smallest spread of data? What is that spread?

1. 哪个四分位的数据分布最集中?该分布幅度是多少?

2. which quarter has the largest spread of data? What is that spread?

2. 哪个四分位的数据分布最分散?该分布幅度是多少?

3. find the interquartile range (*IQR*).

3. 求四分位距(*IQR*)。

4. are there more data in the interval 5–10 or in the interval 10–13? How do you know this?

4. 在区间 5–10 与区间 10–13 中,哪个包含的数据更多?你是如何知道的?

5. which interval has the fewest data in it? How do you know this?

5. 哪个区间包含的数据最少?你是如何知道的?

1. 0–2

1. 0–2

2. 2–4

2. 2–4

3. 10–12

3. 10–12

4. 12–13

4. 12–13

5. need more information

5. 需要更多信息

85.

85.

The following box plot shows the U.S. population for 1990, the latest available year.

以下箱线图展示了美国 1990 年(可获取的最新年份)的人口情况。

1. Are there fewer or more children (age 17 and under) than senior citizens (age 65 and over)? How do you know?

1. 儿童(17 岁及以下)比老年人(65 岁及以上)更少还是更多?你是如何知道的?

2. 12.6% are age 65 and over. Approximately what percentage of the population are working age adults (above age 17 to age 65)?

2. 12.6% 的人为 65 岁及以上。处于工作年龄的成年人(17 岁以上至 65 岁)约占人口百分之多少?

2.4 Box Plots 2.4 箱线图

86\.

86.

In a survey of 20-year-olds in China, Germany, and the United States, people were asked the number of foreign countries they had visited in their lifetime. The following box plots display the results.

在一项针对中国、德国和美国 20 岁人群的调查中,受访者被问及一生中到访过的国家数量。以下箱线图展示了调查结果。

1. In complete sentences, describe what the shape of each box plot implies about the distribution of the data collected.

1. 用完整的句子,描述每个箱线图的形状对所收集数据分布的含义。

2. Have more Americans or more Germans surveyed been to over eight foreign countries?

2. 受访的美国人还是德国人中,到访过八个以上国家的人更多?

3. Compare the three box plots. What do they imply about the foreign travel of 20-year-old residents of the three countries when compared to each other?

3. 比较这三张箱线图。它们对这三个国家 20 岁居民出国旅行情况的相互比较说明了什么?

87.

87.

Given the following box plot, answer the questions.

给定如下箱线图,回答问题。

1. Think of an example (in words) where the data might fit into the above box plot. In 2–5 sentences, write down the example.

1. 想一个(用文字描述的)例子,其数据可能符合上述箱线图。用 2–5 句话写下该例子。

2. What does it mean to have the first and second quartiles so close together, while the second to third quartiles are far apart?

2. 第一与第二四分位数如此接近,而第二与第三四分位数却相距甚远,这意味着什么?

88\.

88.

Given the following box plots, answer the questions.

给定如下箱线图,回答问题。

1. In complete sentences, explain why each statement is false.

1. 用完整的句子,解释下列说法为何错误。

1. Data 1 has more data values above two than Data 2 has above two.

1. 数据 1 中大于 2 的数据值多于 数据 2 中大于 2 的数据值。

2. The data sets cannot have the same mode.

2. 这两个数据集不可能具有相同的众数。

3. For Data 1, there are more data values below four than there are above four.

3. 对于 数据 1,小于 4 的数据值多于大于 4 的数据值。

2. For which group, Data 1 or Data 2, is the value of “7” more likely to be an outlier? Explain why in complete sentences.

2. 对于数据 1 还是数据 2,数值"7"更可能是一个离群值?用完整的句子解释原因。

89.

89.

A survey was conducted of 130 purchasers of new BMW 3 series cars, 130 purchasers of new BMW 5 series cars, and 130 purchasers of new BMW 7 series cars. In it, people were asked the age they were when they purchased their car. The following box plots display the results.

对 130 名新 BMW 3 系、130 名新 BMW 5 系和 130 名新 BMW 7 系的购买者进行了一项调查,询问他们购车时的年龄。以下箱线图展示了调查结果。

1. In complete sentences, describe what the shape of each box plot implies about the distribution of the data collected for that car series.

1. 用完整的句子,描述每个箱线图的形状对所收集的各车系数据分布的含义。

2. Which group is most likely to have an outlier? Explain how you determined that.

2. 哪一组最可能出现离群值?解释你是如何判断的。

3. Compare the three box plots. What do they imply about the age of purchasing a BMW from the series when compared to each other?

3. 比较这三张箱线图。它们对购买各车系 BMW 时的年龄相互比较说明了什么?

4. Look at the BMW 5 series. Which quarter has the smallest spread of data? What is the spread?

4. 看 BMW 5 系。哪个四分位的数据分布最集中?该分布幅度是多少?

5. Look at the BMW 5 series. Which quarter has the largest spread of data? What is that spread?

5. 看 BMW 5 系。哪个四分位的数据分布最分散?该分布幅度是多少?

6. Look at the BMW 5 series. Estimate the interquartile range (IQR).

6. 看 BMW 5 系。估计四分位距(IQR)。

7. Look at the BMW 5 series. Are there more data in the interval 31 to 38 or in the interval 45 to 55? How do you know this?

7. 看 BMW 5 系。在区间 31 到 38 与区间 45 到 55 中,哪个包含的数据更多?你是如何知道的?

8. Look at the BMW 5 series. Which interval has the fewest data in it? How do you know this?

8. 看 BMW 5 系。哪个区间包含的数据最少?你是如何知道的?

1. 31–35

1. 31–35

2. 38–41

2. 38–41

3. 41–64

3. 41–64

90\.

90.

Twenty-five randomly selected students were asked the number of movies they watched the previous week. The results are as follows:

随机选取的 25 名学生被问及他们在上一周观看的电影数量。结果如下:
Table 2.70
\# of moviesFrequency
05
19
26
34
41
表 2.70
电影数量频数
05
19
26
34
41

Construct a box plot of the data.

根据数据绘制箱线图。

2.5 Measures of the Center of the Data 2.5 数据的中心度量

91\.

91.

The most obese countries in the world have obesity rates that range from 11.4% to 74.6%. This data is summarized in the following table.

世界上肥胖率最高的国家的肥胖率介于 11.4% 到 74.6% 之间。该数据汇总于下表。
Table 2.71
Percent of Population ObeseNumber of Countries
11.4–20.4529
20.45–29.4513
29.45–38.454
38.45–47.450
47.45–56.452
56.45–65.451
65.45–74.450
74.45–83.451
表 2.71
肥胖人口百分比国家数量
11.4–20.4529
20.45–29.4513
29.45–38.454
38.45–47.450
47.45–56.452
56.45–65.451
65.45–74.450
74.45–83.451

1. What is the best estimate of the average obesity percentage for these countries?

1. 这些国家肥胖率平均值的最佳估计是多少?

2. The United States has an average obesity rate of 33.9%. Is this rate above average or below?

2. 美国的平均肥胖率为 33.9%。这一比率是高于还是低于平均值?

3. How does the United States compare to other countries?

3. 美国与其他国家相比如何?

92.

92.

Table 2.72 gives the percent of children under five considered to be underweight. What is the best estimate for the mean percentage of underweight children?

表 2.72 给出了被视为体重不足的五岁以下儿童的百分比。体重不足儿童百分比均值的最佳估计是多少?
Table 2.72
Percent of Underweight ChildrenNumber of Countries
16–21.4523
21.45–26.94
26.9–32.359
32.35–37.87
37.8–43.256
43.25–48.71
表 2.72
体重不足儿童百分比国家数量
16–21.4523
21.45–26.94
26.9–32.359
32.35–37.87
37.8–43.256
43.25–48.71

2.6 Skewness and the Mean, Median, and Mode 2.6 偏度与均值、中位数和众数

93\.

93.

The median age of the U.S. population in 1980 was 30.0 years. In 1991, the median age was 33.1 years.

1980 年美国人口的年龄中位数为 30.0 岁。1991 年,年龄中位数为 33.1 岁。

1. What does it mean for the median age to rise?

1. 年龄中位数上升意味着什么?

2. Give two reasons why the median age could rise.

2. 给出年龄中位数可能上升的两条理由。

3. For the median age to rise, is the actual number of children less in 1991 than it was in 1980? Why or why not?

3. 要使年龄中位数上升,1991 年的儿童实际数量是否少于 1980 年?为什么(或为什么不)?

2.7 Measures of the Spread of the Data 2.7 数据的离散程度度量

*Use the following information to answer the next nine exercises:* The population parameters below describe the full-time equivalent number of students (FTES) each year at Lake Tahoe Community College from 1976–1977 through 2004–2005.

*利用以下信息回答接下来九道习题:* 下列总体参数描述了 1976–1977 至 2004–2005 年间,塔霍湖社区学院每年全日制等价学生数(FTES)。

94. A sample of 11 years is taken. About how many are expected to have a FTES of 1014 or above? Explain how you determined your answer.

94. 抽取了 11 年作为样本。预计大约有多少年的 FTES 达到或超过 1014?请解释你是如何得出答案的。

95\. 75% of all years have an FTES:

95. 在所有年份中,75% 的年份其 FTES:

1. at or below: \_\_\_\_\_

1. 在……或以下:\_\_\_\_\_

2. at or above: \_\_\_\_\_

2. 在……或以上:\_\_\_\_\_

96. The population standard deviation = \_\_\_\_\_

96. 总体标准差 = \_\_\_\_\_

97\. What percent of the FTES were from 528.5 to 1447.5? How do you know?

97. 有多少百分比的 FTES 落在 528.5 到 1447.5 之间?你是如何知道的?

98. What is the *IQR*? What does the *IQR* represent?

98. *IQR* 是多少?*IQR* 代表什么?

99\. How many standard deviations away from the mean is the median?

99. 中位数距离均值有多少个标准差?

*Additional Information:* The population FTES for 2005–2006 through 2010–2011 was given in an updated report. The data are reported here.

*补充信息:* 2005–2006 至 2010–2011 年的总体 FTES 在一份更新报告中给出,数据如下。
Table 2.73
Year2005–062006–072007–082008–092009–102010–11
Total FTES1,5851,6901,7351,9352,0211,890
表 2.73
年份2005–062006–072007–082008–092009–102010–11
FTES 总数1,5851,6901,7351,9352,0211,890

100. Calculate the mean, median, standard deviation, the first quartile, the third quartile and the *IQR*. Round to one decimal place.

100. 计算均值、中位数、标准差、第一四分位数、第三四分位数和 *IQR*。保留一位小数。

101\. What additional information is needed to construct a box plot for the FTES for 2005-2006 through 2010-2011 and a box plot for the FTES for 1976-1977 through 2004-2005?

101. 要绘制 2005–2006 至 2010–2011 年 FTES 的箱线图,以及 1976–1977 至 2004–2005 年 FTES 的箱线图,还需要哪些补充信息?

102. Compare the *IQR* for the FTES for 1976–77 through 2004–2005 with the *IQR* for the FTES for 2005-2006 through 2010–2011. Why do you suppose the *IQR*s are so different?

102. 比较 1976–77 至 2004–2005 年 FTES 的 *IQR* 与 2005–2006 至 2010–2011 年 FTES 的 *IQR*。你认为这些 *IQR* 为何差异如此之大?

103\. Three students were applying to the same graduate school. They came from schools with different grading systems. Which student had the best GPA when compared to other students at his school? Explain how you determined your answer.

103. 三名学生申请同一所研究生院,他们来自采用不同评分体系的学校。与各自学校的同学相比,哪名学生的 GPA 最好?请解释你是如何得出答案的。
Table 2.74
StudentGPASchool Average GPASchool Standard Deviation
Thuy2.73.20.8
Vichet877520
Kamala8.680.4
表 2.74
学生GPA学校平均 GPA学校标准差
Thuy2.73.20.8
Vichet877520
Kamala8.680.4

104. A music school has budgeted to purchase three musical instruments. They plan to purchase a piano costing \$3,000, a guitar costing \$550, and a drum set costing \$600. The mean cost for a piano is \$4,000 with a standard deviation of \$2,500. The mean cost for a guitar is \$500 with a standard deviation of \$200. The mean cost for drums is \$700 with a standard deviation of \$100. Which cost is the lowest, when compared to other instruments of the same type? Which cost is the highest when compared to other instruments of the same type. Justify your answer.

104. 一所音乐学校预算购买三件乐器:一架钢琴(3000 美元)、一把吉他(550 美元)和一套鼓(600 美元)。钢琴的平均价格为 4000 美元,标准差为 2500 美元;吉他的平均价格为 500 美元,标准差为 200 美元;鼓的平均价格为 700 美元,标准差为 100 美元。与同类型乐器相比,哪种花费最低?哪种最高?证明你的答案。

105\. An elementary school class ran one mile with a mean of 11 minutes and a standard deviation of three minutes. Rachel, a student in the class, ran one mile in eight minutes. A junior high school class ran one mile with a mean of nine minutes and a standard deviation of two minutes. Kenji, a student in the class, ran 1 mile in 8.5 minutes. A high school class ran one mile with a mean of seven minutes and a standard deviation of four minutes. Nedda, a student in the class, ran one mile in eight minutes.

105. 一所小学班级跑一英里,平均用时 11 分钟,标准差为 3 分钟;该班学生蕾切尔(Rachel)跑一英里用时 8 分钟。一所初中班级跑一英里,平均用时 9 分钟,标准差为 2 分钟;该班学生肯吉(Kenji)跑一英里用时 8.5 分钟。一所高中班级跑一英里,平均用时 7 分钟,标准差为 4 分钟;该班学生内达(Nedda)跑一英里用时 8 分钟。

1. Why is Kenji considered a better runner than Nedda, even though Nedda ran faster than he?

1. 为什么肯吉被认为比内达跑得更好,尽管内达跑得比他快?

2. Who is the fastest runner with respect to his or her class? Explain why.

2. 相对于其所在班级,谁是最快的跑者?解释原因。

106. The most obese countries in the world have obesity rates that range from 11.4% to 74.6%. This data is summarized in Table 14.

106. 世界上肥胖率最高的国家,其肥胖率范围从 11.4% 到 74.6%。该数据总结于表 14。
Table 2.75
Percent of Population ObeseNumber of Countries
11.4–20.4529
20.45–29.4513
29.45–38.454
38.45–47.450
47.45–56.452
56.45–65.451
65.45–74.450
74.45–83.451
表 2.75
肥胖人口百分比国家数量
11.4–20.4529
20.45–29.4513
29.45–38.454
38.45–47.450
47.45–56.452
56.45–65.451
65.45–74.450
74.45–83.451

What is the best estimate of the average obesity percentage for these countries? What is the standard deviation for the listed obesity rates? The United States has an average obesity rate of 33.9%. Is this rate above average or below? How “unusual” is the United States’ obesity rate compared to the average rate? Explain.

这些国家肥胖率平均值的最佳估计是多少?所列肥胖率的标准差是多少?美国的肥胖率平均为 33.9%。这一比率是高于还是低于平均值?与美国的平均比率相比,美国的肥胖率有多"不寻常"?请解释。

107\. Table 2.76 gives the percent of children under five considered to be underweight.

107. 表 2.76 给出了被视为体重不足的 5 岁以下儿童百分比。
Table 2.76
Percent of Underweight ChildrenNumber of Countries
16–21.4523
21.45–26.94
26.9–32.359
32.35–37.87
37.8–43.256
43.25–48.71
表 2.76
体重不足儿童百分比国家数量
16–21.4523
21.45–26.94
26.9–32.359
32.35–37.87
37.8–43.256
43.25–48.71

What is the best estimate for the mean percentage of underweight children? What is the standard deviation? Which interval(s) could be considered unusual? Explain.

体重不足儿童百分比均值的最佳估计是多少?标准差是多少?哪些区间可被视为不寻常?解释。

Bringing It Together: Homework 综合练习:作业

108. Santa Clara County, CA, has approximately 27,873 Japanese-Americans. Their ages are as follows:

108. 加利福尼亚州圣克拉拉县约有 27,873 名日裔美国人。他们的年龄分布如下:
Table 2.77
Age GroupPercent of Community
0–1718.9
18–248.0
25–3422.8
35–4415.0
45–5413.1
55–6411.9
65+10.3
表 2.77
年龄组社区百分比
0–1718.9
18–248.0
25–3422.8
35–4415.0
45–5413.1
55–6411.9
65+10.3

1. Construct a histogram of the Japanese-American community in Santa Clara County, CA. The bars will not be the same width for this example. Why not? What impact does this have on the reliability of the graph?

1. 绘制加利福尼亚州圣克拉拉县日裔美国人社区的直方图。本例中各条形宽度将相同。为什么?这对图形的可靠性有何影响?

2. What percentage of the community is under age 35?

2. 该社区中 35 岁以下的人口占百分之多少?

3. Which box plot most resembles the information above?

3. 哪个箱线图与上述信息最相似?

109\. Javier and Ercilia are supervisors at a shopping mall. Each was given the task of estimating the mean distance that shoppers live from the mall. They each randomly surveyed 100 shoppers. The samples yielded the following information.

109. 哈维尔(Javier)和埃尔西利亚(Ercilia)是一家购物中心的主管。两人各自负责估计购物者居住处到商场的平均距离,各随机调查了 100 名购物者。样本结果如下。
Table 2.78
JavierErcilia
$\overline{x}$6.0 miles6.0 miles
$s$4.0 miles7.0 miles
表 2.78
哈维尔埃尔西利亚
$\overline{x}$6.0 英里6.0 英里
$s$4.0 英里7.0 英里

1. How can you determine which survey was correct ?

1. 你如何判断哪次调查是正确的?

2. Explain what the difference in the results of the surveys implies about the data.

2. 解释两次调查结果的差异对数据意味着什么。

3. If the two histograms depict the distribution of values for each supervisor, which one depicts Ercilia's sample? How do you know?

3. 如果两个直方图分别描述每位主管的数值分布,哪个描述的是埃尔西利亚的样本?你是如何知道的?

4. If the two box plots depict the distribution of values for each supervisor, which one depicts Ercilia’s sample? How do you know?

4. 如果两个箱线图分别描述每位主管的数值分布,哪个描述的是埃尔西利亚的样本?你是如何知道的?

*Use the following information to answer the next three exercises*: We are interested in the number of years students in a particular elementary statistics class have lived in California. The information in the following table is from the entire section.

*利用以下信息回答接下来三道习题:* 我们关注某初级统计班学生在加利福尼亚州居住的年数。下表的资料来自整个班级。
Table 2.79
Number of yearsFrequencyNumber of yearsFrequency
71221
143231
151261
181402
194422
203
Total = 20
表 2.79
年数频数年数频数
71221
143231
151261
181402
194422
203
Total = 20

110. What is the *IQR*?

110. *IQR* 是多少?

1. 8

1. 8

2. 11

2. 11

3. 15

3. 15

4. 35

4. 35

111\. What is the mode?

111. 众数是多少?

1. 19

1. 19

2. 19.5

2. 19.5

3. 14 and 20

3. 14 和 20

4. 22.65

4. 22.65

112. Is this a sample or the entire population?

112. 这是一个样本还是整个总体?

1. sample

1. 样本

2. entire population

2. 整个总体

3. neither

3. 都不是

113. Twenty-five randomly selected students were asked the number of movies they watched the previous week. The results are as follows:

113. 随机询问 25 名学生上周观看的电影部数,结果如下:
Table 2.80
# of moviesFrequency
05
19
26
34
41
表 2.80
电影部数频数
05
19
26
34
41

1. Find the sample mean $\overline{x}$.

1. 求样本均值 $\overline{x}$。

2. Find the approximate sample standard deviation, *s*.

2. 求近似的样本标准差 *s*。

114\. Forty randomly selected students were asked the number of pairs of sneakers they owned. Let *X* = the number of pairs of sneakers owned. The results are as follows:

114. 随机询问 40 名学生拥有运动鞋的鞋数。令 *X* = 拥有的运动鞋鞋数。结果如下:
Table 2.81
*X*Frequency
12
25
38
412
512
60
71
表 2.81
*X*频数
12
25
38
412
512
60
71

1. Find the sample mean $\overset{–}{x}$

1. 求样本均值 $\overset{–}{x}$

2. Find the sample standard deviation, *s*

2. 求样本标准差 *s*

3. Construct a histogram of the data.

3. 绘制数据的直方图。

4. Complete the columns of the chart.

4. 补全图表的各列。

5. Find the first quartile.

5. 求第一四分位数。

6. Find the median.

6. 求中位数。

7. Find the third quartile.

7. 求第三四分位数。

8. Construct a box plot of the data.

8. 绘制数据的箱线图。

9. What percent of the students owned at least five pairs?

9. 拥有至少 5 双的学生占百分之多少?

10. Find the 40th percentile.

10. 求第 40th 百分位数。

11. Find the 90th percentile.

11. 求第 90th 百分位数。

12. Construct a line graph of the data

12. 绘制数据的线图

13. Construct a stemplot of the data

13. 绘制数据的茎叶图

115. Following are the published weights (in pounds) of all of the team members of the San Francisco 49ers from a previous year.

115. 以下是旧金山 49 人队某年前所有队员公布的体重(磅)。

177; 205; 210; 210; 232; 205; 185; 185; 178; 210; 206; 212; 184; 174; 185; 242; 188; 212; 215; 247; 241; 223; 220; 260; 245; 259; 278; 270; 280; 295; 275; 285; 290; 272; 273; 280; 285; 286; 200; 215; 185; 230; 250; 241; 190; 260; 250; 302; 265; 290; 276; 228; 265

177; 205; 210; 210; 232; 205; 185; 185; 178; 210; 206; 212; 184; 174; 185; 242; 188; 212; 215; 247; 241; 223; 220; 260; 245; 259; 278; 270; 280; 295; 275; 285; 290; 272; 273; 280; 285; 286; 200; 215; 185; 230; 250; 241; 190; 260; 250; 302; 265; 290; 276; 228; 265

1. Organize the data from smallest to largest value.

1. 将数据按从小到大排序。

2. Find the median.

2. 求中位数。

3. Find the first quartile.

3. 求第一四分位数。

4. Find the third quartile.

4. 求第三四分位数。

5. Construct a box plot of the data.

5. 绘制数据的箱线图。

6. The middle 50% of the weights are from \_\_\_\_\_\_\_ to \_\_\_\_\_\_\_.

6. 中间的 50% 体重介于 \_\_\_\_\_\_\_ 到 \_\_\_\_\_\_\_ 之间。

7. If our population were all professional football players, would the above data be a sample of weights or the population of weights? Why?

7. 若我们的总体为所有职业橄榄球运动员,上述数据应是体重的样本还是体重的总体?为什么?

8. Assume the population was the San Francisco 49ers. Find:

8. 假设总体为旧金山 49 人队。求:

1. the population mean, *μ*.

1. 总体均值 *μ*。

2. the population standard deviation, *σ*.

2. 总体标准差 *σ*。

3. the weight that is two standard deviations below the mean.

3. 低于均值两个标准差的体重。

4. When Steve Young, quarterback, played football, he weighed 205 pounds. How many standard deviations above or below the mean was he?

4. 四分卫史蒂夫·扬(Steve Young)打橄榄球时体重为 205 磅。他比均值高或低多少个标准差?

9. That same year, the mean weight for the Dallas Cowboys was 240.08 pounds with a standard deviation of 44.38 pounds. Emmit Smith weighed in at 209 pounds. With respect to his team, who was lighter, Smith or Young? How did you determine your answer?

9. 同一年,达拉斯牛仔队的平均体重为 240.08 磅,标准差为 44.38 磅。埃米特·史密斯(Emmit Smith)体重为 209 磅。相对于各自球队,谁更轻,史密斯还是扬?你是如何得出答案的?

116\. One hundred teachers attended a seminar on mathematical problem solving. The attitudes of a representative sample of 12 of the teachers were measured before and after the seminar. A positive number for change in attitude indicates that a teacher's attitude toward math became more positive. The 12 change scores are as follows:

116. 100 名教师参加了一个数学解题研讨会。在研讨会前后测量了其中 12 名教师(代表性样本)的态度。态度变化的正值表示该教师对数学的态度变得更正面的。这 12 个变化分数如下:

3; 8; –1; 2; 0; 5; –3; 1; –1; 6; 5; –2

3; 8; –1; 2; 0; 5; –3; 1; –1; 6; 5; –2

1. What is the mean change score?

1. 变化分数的均值是多少?

2. What is the standard deviation for this population?

2. 该总体的标准差是多少?

3. What is the median change score?

3. 变化分数的中位数是多少?

4. Find the change score that is 2.2 standard deviations below the mean.

4. 求低于均值 2.2 个标准差的变化分数。

117. Refer to Figure 2.50 determine which of the following are true and which are false. Explain your solution to each part in complete sentences.

117. 参考图 2.50,判断下列关于哪个为真、哪个为假。请用完整的句子解释每一部分的解答。

1. The medians for all three graphs are the same.

1. 三幅图的中位数相同。

2. We cannot determine if any of the means for the three graphs is different.

2. 我们无法判断三幅图中是否有均值不同。

3. The standard deviation for graph b is larger than the standard deviation for graph a.

3. 图 b 的标准差大于图 a 的标准差。

4. We cannot determine if any of the third quartiles for the three graphs is different.

4. 我们无法判断三幅图中是否有第三四分位数不同。

118\. In a recent issue of the IEEE Spectrum, 84 engineering conferences were announced. Four conferences lasted two days. Thirty-six lasted three days. Eighteen lasted four days. Nineteen lasted five days. Four lasted six days. One lasted seven days. One lasted eight days. One lasted nine days. Let *X* = the length (in days) of an engineering conference.

118. 在最近一期的 IEEE Spectrum 上公布了 84 个工程会议。其中 4 个会议持续 2 天,36 个持续 3 天,18 个持续 4 天,19 个持续 5 天,4 个持续 6 天,1 个持续 7 天,1 个持续 8 天,1 个持续 9 天。令 *X* = 工程会议的时长(天)。

1. Organize the data in a chart.

1. 将数据整理成图表。

2. Find the median, the first quartile, and the third quartile.

2. 求中位数、第一四分位数和第三四分位数。

3. Find the 65th percentile.

3. 求第 65th 百分位数。

4. Find the 10th percentile.

4. 求第 10th 百分位数。

5. Construct a box plot of the data.

5. 绘制数据的箱线图。

6. The middle 50% of the conferences last from \_\_\_\_\_\_\_ days to \_\_\_\_\_\_\_ days.

6. 中间的 50% 会议持续 \_\_\_\_\_\_\_ 天到 \_\_\_\_\_\_\_ 天。

7. Calculate the sample mean of days of engineering conferences.

7. 计算工程会议天数的样本均值。

8. Calculate the sample standard deviation of days of engineering conferences.

8. 计算工程会议天数的样本标准差。

9. Find the mode.

9. 求众数。

10. If you were planning an engineering conference, which would you choose as the length of the conference: mean; median; or mode? Explain why you made that choice.

10. 如果你要筹划一个工程会议,你会选择均值、中位数还是众数作为会议时长?解释你的选择理由。

11. Give two reasons why you think that three to five days seem to be popular lengths of engineering conferences.

11. 给出两个理由,说明为什么你认为三到五天似乎是工程会议常见的时长。

119. A survey of enrollment at 35 community colleges across the United States yielded the following figures:

119. 一项针对美国 35 所社区学院入学人数的调查得到以下数字:

6414; 1550; 2109; 9350; 21828; 4300; 5944; 5722; 2825; 2044; 5481; 5200; 5853; 2750; 10012; 6357; 27000; 9414; 7681; 3200; 17500; 9200; 7380; 18314; 6557; 13713; 17768; 7493; 2771; 2861; 1263; 7285; 28165; 5080; 11622

6414; 1550; 2109; 9350; 21828; 4300; 5944; 5722; 2825; 2044; 5481; 5200; 5853; 2750; 10012; 6357; 27000; 9414; 7681; 3200; 17500; 9200; 7380; 18314; 6557; 13713; 17768; 7493; 2771; 2861; 1263; 7285; 28165; 5080; 11622

1. Organize the data into a chart with five intervals of equal width. Label the two columns "Enrollment" and "Frequency."

1. 将数据整理成图表,分为 5 个等宽区间。将两列分别标注为"入学人数"和"频数"。

2. Construct a histogram of the data.

2. 绘制数据的直方图。

3. If you were to build a new community college, which piece of information would be more valuable: the mode or the mean?

3. 如果你要新建一所社区学院,哪条信息更有价值:众数还是均值?

4. Calculate the sample mean.

4. 计算样本均值。

5. Calculate the sample standard deviation.

5. 计算样本标准差。

6. A school with an enrollment of 8000 would be how many standard deviations away from the mean?

6. 一所入学人数为 8000 的学校距离均值有多少个标准差?

*Use the following information to answer the next two exercises.* *X* = the number of days per week that 100 clients use a particular exercise facility.

*利用以下信息回答接下来两道习题。* *X* = 100 名客户每周使用某健身设施的天数。
Table 2.82
*x*Frequency
03
112
233
328
411
59
64
表 2.82
*x*频数
03
112
233
328
411
59
64

120. The 80th percentile is \_\_\_\_\_

120. 第 80th 百分位数是 \_\_\_\_\_

1. 5

1. 5

2. 80

2. 80

3. 3

3. 3

4. 4

4. 4

121. The number that is 1.5 standard deviations BELOW the mean is approximately \_\_\_\_\_

121. 低于均值 1.5 个标准差的数约为 \_\_\_\_\_

1. 0.7

1. 0.7

2. 4.8

2. 4.8

3. –2.8

3. –2.8

4. Cannot be determined

4. 无法确定

122\. Suppose that a publisher conducted a survey asking adult consumers the number of fiction paperback books they had purchased in the previous month. The results are summarized in the Table 2.83.

122. 假设一家出版商进行了一项调查,询问成年消费者上月购买的虚构平装书数量。结果总结于表 2.83。
Table 2.83
# of booksFreq.Rel. Freq.
018
124
224
322
415
510
75
91
表 2.83
图书部数频数相对频数
018
124
224
322
415
510
75
91

1. Are there any outliers in the data? Use an appropriate numerical test involving the *IQR* to identify outliers, if any, and clearly state your conclusion.

1. 数据中是否存在离群值?使用涉及 *IQR* 的适当数值检验来识别离群值(如有),并清楚陈述你的结论。

2. If a data value is identified as an outlier, what should be done about it?

2. 若一个数据值被识别为离群值,应如何处理?

3. Are any data values further than two standard deviations away from the mean? In some situations, statisticians may use this criteria to identify data values that are unusual, compared to the other data values. (Note that this criteria is most appropriate to use for data that is mound-shaped and symmetric, rather than for skewed data.)

3. 是否有任何数据值距离均值超过两个标准差?在某些情况下,统计学家可能用此标准来识别相对于其他数据值而言不寻常的数据值。(注意,该标准最适用于呈钟形且对称的数据,而非偏斜数据。)

4. Do parts a and c of this problem give the same answer?

4. 本题的 a 部分与 c 部分给出的答案是否相同?

5. Examine the shape of the data. Which part, a or c, of this question gives a more appropriate result for this data?

5. 考察数据的形状。本题的 a 部分还是 c 部分对此数据给出更合适的结果?

6. Based on the shape of the data which is the most appropriate measure of center for this data: mean, median or mode?

6. 基于数据的形状,对此数据最适当的集中趋势度量是哪个:均值、中位数还是众数?