← 学习库 Introductory Statistics (OpenStax) · 中英对照 目录

1 Sampling and Data 抽样与数据

本页译自 OpenStax《Introductory Statistics》第 1 章 Sampling and Data。公式经本地 MathJax 渲染,自定义宏已注入。

Introduction 引言

You are probably asking yourself the question, "When and where will I use statistics?" If you read any newspaper, watch television, or use the Internet, you will see statistical information. There are statistics about crime, sports, education, politics, and real estate. Typically, when you read a newspaper article or watch a television news program, you are given sample information. With this information, you may make a decision about the correctness of a statement, claim, or "fact." Statistical methods can help you make the "best educated guess."

你大概正在问自己这样一个问题:"我会在什么时候、什么地方用到统计学?"如果你读报纸、看电视或使用互联网,都会接触到统计信息。有关犯罪、体育、教育、政治和房地产的统计信息随处可见。通常,当你读一篇报纸文章或观看电视新闻节目时,你得到的是样本信息。借助这些信息,你可以判断某个陈述、主张或"事实"是否正确。统计方法能帮助你做出"最有依据的判断"。

Since you will undoubtedly be given statistical information at some point in your life, you need to know some techniques for analyzing the information thoughtfully. Think about buying a house or managing a budget. Think about your chosen profession. The fields of economics, business, psychology, education, biology, law, computer science, police science, and early childhood development require at least one course in statistics.

既然在你人生的某个时刻无疑会接触到统计信息,你就需要掌握一些审慎分析这些信息的方法。想想买房或管理预算,想想你所选择的职业。经济学、商学、心理学、教育学、生物学、法学、计算机科学、警政科学以及儿童早期发展等领域,都至少需要修读一门统计学课程。

Included in this chapter are the basic ideas and words of probability and statistics. You will soon understand that statistics and probability work together. You will also learn how data are gathered and what "good" data can be distinguished from "bad."

本章包含了概率与统计学的基本概念和术语。你很快就会明白,统计学与概率是协同工作的。你还将学习数据是如何收集的,以及如何把"好"数据与"坏"数据区分开来。

1.1 Definitions of Statistics, Probability, and Key Terms 1.1 统计学、概率与关键术语的定义

The science of statistics deals with the collection, analysis, interpretation, and presentation of data. We see and use data in our everyday lives.

统计学的科学涉及数据的收集、分析、解释与呈现。我们在日常生活中看到并使用数据。

In your classroom, try this exercise. Have class members write down the average time (in hours, to the nearest half-hour) they sleep per night. Your instructor will record the data. Then create a simple graph (called a dot plot) of the data. A dot plot consists of a number line and dots (or points) positioned above the number line. For example, consider the following data:

在你的课堂上,试着做这个练习。让班级成员写下他们每晚睡眠的平均时间(以小时计,精确到最近的半小时)。你的老师会记录这些数据。然后为这些数据制作一个简单的图形(称为点图,dot plot)。点图由一条数轴以及位于数轴上方的圆点(或点)组成。例如,考虑以下数据:

5; 5.5; 6; 6; 6; 6.5; 6.5; 6.5; 6.5; 7; 7; 8; 8; 9

5;5.5;6;6;6;6.5;6.5;6.5;6.5;7;7;8;8;9

The dot plot for this data would be as follows:

这组数据的点图如下所示:

Does your dot plot look the same as or different from the example? Why? If you did the same example in an English class with the same number of students, do you think the results would be the same? Why or why not?

你的点图与示例看起来相同还是不同?为什么?如果你在一个英语课上用相同数量的学生做同样的例子,你认为结果会相同吗?为什么相同,或为什么不同?

Where do your data appear to cluster? How might you interpret the clustering?

你的数据看起来聚集在何处?你可能会如何解释这种聚集?

The questions above ask you to analyze and interpret your data. With this example, you have begun your study of statistics.

上面的问题要求你分析和解释自己的数据。通过这个例子,你已经开始了统计学的学习。

In this course, you will learn how to organize and summarize data. Organizing and summarizing data is called descriptive statistics. Two ways to summarize data are by graphing and by using numbers (for example, finding an average). After you have studied probability and probability distributions, you will use formal methods for drawing conclusions from "good" data. The formal methods are called inferential statistics. Statistical inference uses probability to determine how confident we can be that our conclusions are correct.

在本课程中,你将学习如何整理和概括数据。整理和概括数据称为描述统计学。概括数据有两种方式:作图和使用数值(例如计算平均数)。在学习了概率与概率分布之后,你将使用正规的方法从"好"数据中得出结论。这些正规方法称为推断统计学。统计推断利用概率来确定我们对结论正确性的把握程度。

Effective interpretation of data (inference) is based on good procedures for producing data and thoughtful examination of the data. You will encounter what will seem to be too many mathematical formulas for interpreting data. The goal of statistics is not to perform numerous calculations using the formulas, but to gain an understanding of your data. The calculations can be done using a calculator or a computer. The understanding must come from you. If you can thoroughly grasp the basics of statistics, you can be more confident in the decisions you make in life.

对数据的有效解释(推断)建立在良好的数据产生程序和对数据的审慎考察之上。你会遇到看似过多的、用于解释数据的数学公式。统计学的目标不是用这些公式做大量计算,而是去理解你的数据。计算可以用计算器或计算机完成,而理解必须来自你自己。如果你能透彻掌握统计学的基础,你就能在生活中做出更自信的决策。

Probability 概率

Probability is a mathematical tool used to study randomness. It deals with the chance (the likelihood) of an event occurring. For example, if you toss a fair coin four times, the outcomes may not be two heads and two tails. However, if you toss the same coin 4,000 times, the outcomes will be close to half heads and half tails. The expected theoretical probability of heads in any one toss is $\frac{1}{2}$ or 0.5. Even though the outcomes of a few repetitions are uncertain, there is a regular pattern of outcomes when there are many repetitions. After reading about the English statistician Karl Pearson who tossed a coin 24,000 times with a result of 12,012 heads, one of the authors tossed a coin 2,000 times. The results were 996 heads. The fraction $\frac{996}{2000}$ is equal to 0.498 which is very close to 0.5, the expected probability.

概率是一种用于研究随机性的数学工具。它涉及某个事件发生的可能性(likelihood)。例如,如果你把一枚均匀硬币抛掷四次,结果未必是两个正面和两个反面。然而,如果你把同一枚硬币抛掷 4000 次,结果将接近一半正面、一半反面。任意一次抛掷出现正面的期望理论概率是 $\frac{1}{2}$,即 0.5。尽管少数几次重复的结果是不可预测的,但当重复次数很多时,结果会呈现出一种规律性的模式。在了解到英国统计学家卡尔·皮尔逊(Karl Pearson)抛掷硬币 24000 次、得到 12012 次正面的结果之后,本书的一位作者抛掷了一枚硬币 2000 次,结果为 996 次正面。分数 $\frac{996}{2000}$ 等于 0.498,非常接近期望值 0.5。

The theory of probability began with the study of games of chance such as poker. Predictions take the form of probabilities. To predict the likelihood of an earthquake, of rain, or whether you will get an A in this course, we use probabilities. Doctors use probability to determine the chance of a vaccination causing the disease the vaccination is supposed to prevent. A stockbroker uses probability to determine the rate of return on a client's investments. You might use probability to decide to buy a lottery ticket or not. In your study of statistics, you will use the power of mathematics through probability calculations to analyze and interpret your data.

概率论始于对扑克等机会游戏的研究。预测以概率的形式出现。为了预测地震、降雨发生的可能性,或者你在本课程中能否拿到 A,我们会使用概率。医生利用概率来确定疫苗引发本应由疫苗预防的那种疾病的概率。股票经纪人利用概率来确定客户投资的回报率。你可能会用概率来决定是否购买彩票。在统计学的学习中,你将通过概率计算运用数学的力量来分析和解释你的数据。

Key Terms 关键术语

In statistics, we generally want to study a population. You can think of a population as a collection of persons, things, or objects under study. To study the population, we select a sample. The idea of sampling is to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. Data are the result of sampling from a population.

在统计学中,我们通常希望研究一个总体。你可以把总体看作处于研究之中的人、物或对象的集合。为了研究总体,我们选取一个样本。抽样的思想是选取较大总体的一部分(或子集)并研究这一部分(即样本),从而获得关于总体的信息。数据是从总体中抽样得到的结果。

Because it takes a lot of time and money to examine an entire population, sampling is a very practical technique. If you wished to compute the overall grade point average at your school, it would make sense to select a sample of students who attend the school. The data collected from the sample would be the students' grade point averages. In presidential elections, opinion poll samples of 1,000–2,000 people are taken. The opinion poll is supposed to represent the views of the people in the entire country. Manufacturers of canned carbonated drinks take samples to determine if a 16 ounce can contains 16 ounces of carbonated drink.

由于考察整个总体需要大量的时间和金钱,抽样是一种非常实用的技术。如果你想计算你所在学校的整体平均绩点(GPA),选取在校学生的一个样本是合理的。从样本收集到的数据就是这些学生的平均绩点。在总统选举中,会抽取 1000–2000 人的民意调查样本。民意调查旨在代表全国民众的观点。罐装碳酸饮料的制造商会抽取样本,以确定一罐 16 盎司的饮料是否含有 16 盎司的碳酸饮料。

From the sample data, we can calculate a statistic. A statistic is a number that represents a property of the sample. For example, if we consider one math class to be a sample of the population of all math classes, then the average number of points earned by students in that one math class at the end of the term is an example of a statistic. The statistic is an estimate of a population parameter. A parameter is a numerical characteristic of the whole population that can be estimated by a statistic. Since we considered all math classes to be the population, then the average number of points earned per student over all the math classes is an example of a parameter.

根据样本数据,我们可以计算出一个统计量。统计量是一个代表样本某种特征的数字。例如,如果我们把某一个数学班看作所有数学班这一总体的一个样本,那么该班学生在学期末获得的平均分数就是一个统计量。这个统计量是对总体参数的一个估计。参数(parameter)是总体的数值特征,可以由统计量来估计。既然我们把所有数学班看作总体,那么所有数学班每个学生平均获得的分数就是一个参数。

One of the main concerns in the field of statistics is how accurately a statistic estimates a parameter. The accuracy really depends on how well the sample represents the population. The sample must contain the characteristics of the population in order to be a representative sample. We are interested in both the sample statistic and the population parameter in inferential statistics. In a later chapter, we will use the sample statistic to test the validity of the established population parameter.

统计学领域的主要关注点之一,是统计量对参数的估计有多准确。这种准确性实际上取决于样本对总体的代表程度。样本必须包含总体的特征,才能成为一个有代表性的样本。在推断统计学中,我们既关注样本统计量,也关注总体参数。在后面的章节中,我们将用样本统计量来检验已确立的总体参数的有效性。

A variable, usually notated by capital letters such as *X* and *Y*, is a characteristic or measurement that can be determined for each member of a population. Variables may be numerical or categorical. Numerical variables take on values with equal units such as weight in pounds and time in hours. Categorical variables place the person or thing into a category. If we let *X* equal the number of points earned by one math student at the end of a term, then *X* is a numerical variable. If we let *Y* be a person's party affiliation, then some examples of *Y* include Republican, Democrat, and Independent. *Y* is a categorical variable. We could do some math with values of *X* (calculate the average number of points earned, for example), but it makes no sense to do math with values of *Y* (calculating an average party affiliation makes no sense).

变量(variable),通常用大写字母如 *X* 和 *Y* 表示,是对总体中每个成员都能确定的特征或度量。变量可以是数值型的或分类的。数值型变量取具有相等单位的数值,例如以磅计的重量和以小时计的时间。分类变量把人或物归入某一类别。如果我们令 *X* 等于一名数学专业学生在学期末获得的分数,那么 *X* 是一个数值型变量。如果我们令 *Y* 表示一个人的政党归属,那么 *Y* 的一些例子包括共和党、民主党和无党派。*Y* 是一个分类变量。我们可以对 *X* 的取值做数学运算(例如计算平均得分),但对 *Y* 的取值做数学运算没有意义(计算平均政党归属没有意义)。

Data are the actual values of the variable. They may be numbers or they may be words. Datum is a single value.

数据是变量的实际取值。它们可以是数字,也可以是文字。Datum(数据点)是单个取值。

Two words that come up often in statistics are mean and proportion. If you were to take three exams in your math classes and obtain scores of 86, 75, and 92, you would calculate your mean score by adding the three exam scores and dividing by three (your mean score would be 84.3 to one decimal place). If, in your math class, there are 40 students and 22 are men and 18 are women, then the proportion of men students is $\frac{22}{40}$ and the proportion of women students is $\frac{18}{40}$. Mean and proportion are discussed in more detail in later chapters.

统计学中经常出现的两个词是均值(mean)和比例(proportion)。如果你在数学课上参加三次考试,成绩分别为 86、75 和 92,你会把这三个考试分数相加再除以三来计算你的平均分数(你的平均分保留一位小数为 84.3)。如果在你的数学班上有 40 名学生,其中 22 名为男生、18 名为女生,那么男生学生的比例是 $\frac{22}{40}$,女生学生的比例是 $\frac{18}{40}$。均值和比例将在后面的章节中更详细地讨论。

The words "mean" and "average" are often used interchangeably. The substitution of one word for the other is common practice. The technical term is "arithmetic mean," and "average" is technically a center location. However, in practice among non-statisticians, "average" is commonly accepted for "arithmetic mean."

"mean"(均值)和"average"(平均数)这两个词经常互换使用。用一个词替代另一个词是常见做法。技术术语是"算术平均数"(arithmetic mean),而"average"严格来说是一种中心位置。然而,在非统计学家群体的实践中,"average"通常被接受为"算术平均数"的同义用法。

Problem 问题

Determine what the key terms refer to in the following study. We want to know the average (mean) amount of money first year college students spend at ABC College on school supplies that do not include books. We randomly surveyed 100 first year students at the college. Three of those students spent \$150, \$200, and \$225, respectively.

确定下列研究中各关键术语所指代的对象。我们想知道 ABC 学院大一学生在不含书籍的学校用品上花费的平均(均值)金额。我们随机调查了该学院的 100 名大一学生。其中三名学生分别花费了 150 美元、200 美元和 225 美元。

Solution 解答

The population is all first year students attending ABC College this term.

总体是本学期在 ABC 学院就读的所有大一学生。

The sample could be all students enrolled in one section of a beginning statistics course at ABC College (although this sample may not represent the entire population).

样本可以是 ABC 学院某入门统计学课程一个教学班中的所有学生(尽管这个样本可能无法代表整个总体)。

The parameter is the average (mean) amount of money spent (excluding books) by first year college students at ABC College this term.

参数是本学期 ABC 学院大一学生在(不含书籍的)学校用品上花费的平均(均值)金额。

The statistic is the average (mean) amount of money spent (excluding books) by first year college students in the sample.

统计量是样本中大一学生在(不含书籍的)学校用品上花费的平均(均值)金额。

The variable could be the amount of money spent (excluding books) by one first year student. Let *X* = the amount of money spent (excluding books) by one first year student attending ABC College.

变量可以是一名大一学生花费的(不含书籍的)金额。令 *X* = 一名就读于 ABC 学院的大一学生花费的(不含书籍的)金额。

The data are the dollar amounts spent by the first year students. Examples of the data are \$150, \$200, and \$225.

数据是大一学生花费的美元金额。数据的例子有 150 美元、200 美元和 225 美元。

Determine what the key terms refer to in the following study. We want to know the average (mean) amount of money spent on school uniforms each year by families with children at Knoll Academy. We randomly survey 100 families with children in the school. Three of the families spent \$65, \$75, and \$95, respectively.

确定下列研究中各关键术语所指代的对象。我们想知道 Knoll 学院有子女的家庭每年在校服上花费的平均(均值)金额。我们随机调查了该学校 100 个有子女的家庭。其中三个家庭分别花费了 65 美元、75 美元和 95 美元。

Problem 问题

Determine what the key terms refer to in the following study.

确定下列研究中各关键术语所指代的对象。

A study was conducted at a local college to analyze the average cumulative GPA’s of students who graduated last year. Fill in the letter of the phrase that best describes each of the items below.

在本地一所学院进行了一项研究,以分析去年毕业学生的平均累计 GPA。请填写最能描述下列各项的短语所对应的字母。

1\. Population\_\_\_\_\_ 2. Statistic \_\_\_\_\_ 3. Parameter \_\_\_\_\_ 4. Sample \_\_\_\_\_ 5. Variable \_\_\_\_\_ 6. Data \_\_\_\_\_

1. 总体_____ 2. 统计量 _____ 3. 参数 _____ 4. 样本 _____ 5. 变量 _____ 6. 数据 _____

1. all students who attended the college last year

1. 去年就读于该学院的所有学生

2. the cumulative GPA of one student who graduated from the college last year

2. 去年从该院毕业的一名大学生的累计 GPA

3. 3.65, 2.80, 1.50, 3.90

3. 3.65、2.80、1.50、3.90

4. a group of students who graduated from the college last year, randomly selected

4. 去年从该院毕业、经随机选取的一组学生

5. the average cumulative GPA of students who graduated from the college last year

5. 去年从该院毕业学生的平均累计 GPA

6. all students who graduated from the college last year

6. 去年从该院毕业的所有学生

7. the average cumulative GPA of students in the study who graduated from the college last year

7. 研究中去年从该院毕业学生的平均累计 GPA

Solution 解答

1\. f; 2. g; 3. e; 4. d; 5. b; 6. c

1. f;2. g;3. e;4. d;5. b;6. c

Problem 问题

Determine what the key terms refer to in the following study.

确定下列研究中各关键术语所指代的对象。

As part of a study designed to test the safety of automobiles, the National Transportation Safety Board collected and reviewed data about the effects of an automobile crash on test dummies. Here is the criterion they used:

作为一项旨在测试汽车安全性的研究的一部分,国家运输安全委员会(National Transportation Safety Board)收集并审查了有关汽车碰撞对测试假人影响的数据。以下是他们所采用的标准:

| | |

| | |

|-----------------------------|--------------------------------------|

|-----------------------------|--------------------------------------|

| Speed at which Cars Crashed | Location of "drivers" (i.e. dummies) |

| 汽车碰撞速度 | "驾驶员"(即假人)的位置 |

| 35 miles/hour | Front Seat |

| 35 英里/小时 | 前座 |

Table 1.1

表 1.1

Cars with dummies in the front seats were crashed into a wall at a speed of 35 miles per hour. We want to know the proportion of dummies in the driver’s seat that would have had head injuries, if they had been actual drivers. We start with a simple random sample of 75 cars.

前座装有假人的汽车以 35 英里/小时的速度撞向一面墙。我们想知道,如果它们是真正的驾驶员,驾驶座上的假人中会发生头部受伤的比例。我们从一个包含 75 辆汽车的简单随机样本开始。

Solution 解答

The population is all cars containing dummies in the front seat.

总体是所有前座装有假人的汽车。

The sample is the 75 cars, selected by a simple random sample.

样本是这 75 辆汽车,通过简单随机抽样选取。

The parameter is the proportion of driver dummies (if they had been real people) who would have suffered head injuries in the population.

参数是总体中驾驶员假人(如果他们是真人)会发生头部受伤的比例。

The statistic is proportion of driver dummies (if they had been real people) who would have suffered head injuries in the sample.

统计量是样本中驾驶员假人(如果他们是真人)会发生头部受伤的比例。

The variable *X* = whether a dummy (if it had been a real person) who would have suffered head injuries.

变量 *X* = 一名假人(如果他是真人)是否会发生头部受伤。

The data are either: yes, had head injury, or no, did not.

数据为以下两者之一:是,发生了头部受伤;或否,没有发生。

Problem 问题

Determine what the key terms refer to in the following study.

确定下列研究中各关键术语所指代的对象。

An insurance company would like to determine the proportion of all medical doctors who have been involved in one or more malpractice lawsuits. The company selects 500 doctors at random from a professional directory and determines the number in the sample who have been involved in a malpractice lawsuit.

一家保险公司希望确定所有内科医生中曾涉及一起或多起医疗事故诉讼的比例。该公司从一本专业名录中随机选取 500 名医生,并确定样本中有多少人曾涉及医疗事故诉讼。

Solution 解答

The population is all medical doctors listed in the professional directory.

总体是专业目录中列出的所有医生。

The parameter is the proportion of medical doctors who have been involved in one or more malpractice suits in the population.

参数是总体中曾涉及一起或多起医疗事故诉讼的医生所占的比例。

The sample is the 500 doctors selected at random from the professional directory.

样本是从专业目录中随机抽取的 500 名医生。

The statistic is the proportion of medical doctors who have been involved in one or more malpractice suits in the sample.

统计量是样本中曾涉及一起或多起医疗事故诉讼的医生所占的比例。

The variable *X* = whether an individual doctor has been involved in a malpractice suit.

变量 *X* = 一位医生是否曾涉及医疗事故诉讼。

The data are either: yes, was involved in one or more malpractice lawsuits, or no, was not.

数据是以下二者之一:是,曾涉及一起或多起医疗事故诉讼;或否,未曾涉及。

Do the following exercise collaboratively with up to four people per group. Find a population, a sample, the parameter, the statistic, a variable, and data for the following study: You want to determine the average (mean) number of glasses of milk college students drink per day. Suppose yesterday, in your English class, you asked five students how many glasses of milk they drank the day before. The answers were 1, 0, 1, 3, and 4 glasses of milk.

以下练习请以小组合作方式完成,每组最多四人。针对以下研究找出总体、样本、参数、统计量、一个变量和数据:你想确定大学生每天平均(均值)喝多少杯牛奶。假设昨天你在英语课上问了五名学生他们前一天喝了几杯牛奶,答案是 1、0、1、3 和 4 杯。

1.2 Data, Sampling, and Variation in Data and Sampling 1.2 数据、抽样与数据及抽样中的变异

Data may come from a population or from a sample. Lowercase letters like $x$ or $y$ generally are used to represent data values. Most data can be put into the following categories:

数据可能来自总体,也可能来自样本。像 $x$ 或 $y$ 这样的小写字母通常表示数据值。大多数数据可归入以下类别:

Qualitative data are the result of categorizing or describing attributes of a population. Qualitative data are also often called categorical data. Hair color, blood type, ethnic group, the car a person drives, and the street a person lives on are examples of qualitative data. Qualitative data are generally described by words or letters. For instance, hair color might be black, dark brown, light brown, blonde, gray, or red. Blood type might be AB+, O-, or B+. Researchers often prefer to use quantitative data over qualitative data because it lends itself more easily to mathematical analysis. For example, it does not make sense to find an average hair color or blood type.

定性数据是对总体的属性进行分类或描述的结果。定性数据也常被称为分类数据。发色、血型、种族、一个人所开的车,以及一个人所居住的街道,都是定性数据的例子。定性数据通常用文字或字母来描述。例如,发色可能是黑色、深棕色、浅棕色、金色、灰色或红色;血型可能是 AB+、O- 或 B+。研究者往往更倾向于使用定量数据而非定性数据,因为它更便于数学分析。例如,求发色或血型的平均值是没有意义的。

Quantitative data are always numbers. Quantitative data are the result of counting or measuring attributes of a population. Amount of money, pulse rate, weight, number of people living in your town, and number of students who take statistics are examples of quantitative data. Quantitative data may be either discrete or continuous.

定量数据总是数值。定量数据是对总体属性进行计数或测量的结果。金额、脉搏率、体重、你所居住城镇的人口数,以及修读统计学课程的学生数,都是定量数据的例子。定量数据可以是离散的,也可以是连续的。

All data that are the result of counting are called quantitative discrete data. These data take on only certain numerical values. If you count the number of phone calls you receive for each day of the week, you might get values such as zero, one, two, or three.

所有由计数得到的数据称为定量离散数据。这些数据只取某些特定的数值。如果你统计一周中每天接到的电话数,可能会得到诸如零、一、二或三这样的数值。

Data that are not only made up of counting numbers, but that may include fractions, decimals, or irrational numbers, are called quantitative continuous data. Continuous data are often the results of measurements like lengths, weights, or times. A list of the lengths in minutes for all the phone calls that you make in a week, with numbers like 2.4, 7.5, or 11.0, would be quantitative continuous data.

不仅由计数整数构成,而且可能包含分数、小数或无理数的数据,称为定量连续数据。连续数据通常是测量(如长度、重量或时间)的结果。一份你一周中所打电话的时长(分钟)清单,其中的数值如 2.4、7.5 或 11.0,就是定量连续数据。

Data Sample of Quantitative Discrete Data 定量离散数据的数据样例

The data are the number of books students carry in their backpacks. You sample five students. Two students carry three books, one student carries four books, one student carries two books, and one student carries one book. The numbers of books (three, four, two, and one) are the quantitative discrete data.

数据是学生们背包中所带书本的数量。你抽取了五名学生。两名学生带三本书,一名学生带四本书,一名学生带两本书,还有一名学生带一本书。这些书本数量(三、四、二、一)就是定量离散数据。

The data are the number of machines in a gym. You sample five gyms. One gym has 12 machines, one gym has 15 machines, one gym has ten machines, one gym has 22 machines, and the other gym has 20 machines. What type of data is this?

数据是健身房中机器的数量。你抽取了五家健身房。一家健身房有 12 台机器,一家有 15 台,一家有 10 台,一家有 22 台,另一家有 20 台。这是什么类型的数据?

Data Sample of Quantitative Continuous Data 定量连续数据的数据样例

The data are the weights of backpacks with books in them. You sample the same five students. The weights (in pounds) of their backpacks are 6.2, 7, 6.8, 9.1, 4.3. Notice that backpacks carrying three books can have different weights. Weights are quantitative continuous data.

数据是装有书的背包的重量。你抽取的是同样的五名学生。他们背包的重量(磅)分别为 6.2、7、6.8、9.1、4.3。注意,装三本书的背包也可能有不同的重量。重量属于定量连续数据。

The data are the areas of lawns in square feet. You sample five houses. The areas of the lawns are 144 sq. feet, 160 sq. feet, 190 sq. feet, 180 sq. feet, and 210 sq. feet. What type of data is this?

数据是草坪的面积(平方英尺)。你抽取了五所房屋。草坪面积分别为 144 平方英尺、160 平方英尺、190 平方英尺、180 平方英尺和 210 平方英尺。这是什么类型的数据?

You go to the supermarket and purchase three cans of soup (19 ounces tomato bisque, 14.1 ounces lentil, and 19 ounces Italian wedding), two packages of nuts (walnuts and peanuts), four different kinds of vegetable (broccoli, cauliflower, spinach, and carrots), and two desserts (16 ounces pistachio ice cream and 32 ounces chocolate chip cookies).

你去超市买了三罐汤(19 盎司番茄浓汤、14.1 盎司扁豆汤和 19 盎司意式婚礼汤)、两包坚果(核桃和花生)、四种不同的蔬菜(西兰花、花椰菜、菠菜和胡萝卜),以及两种甜点(16 盎司开心果冰淇淋和 32 盎司巧克力曲奇)。

Problem 问题

Name data sets that are quantitative discrete, quantitative continuous, and qualitative.

举出分别属于定量离散、定量连续和定性的数据集。

Solution 解答

One Possible Solution:

一种可能的解答:

Try to identify additional data sets in this example.

试着在这个例子里再找出其他数据集。

The data are the colors of backpacks. Again, you sample the same five students. One student has a red backpack, two students have black backpacks, one student has a green backpack, and one student has a gray backpack. The colors red, black, black, green, and gray are qualitative data.

数据是背包的颜色。同样,你抽取这五名学生。一名学生有红色背包,两名学生有黑色背包,一名学生有绿色背包,还有一名学生有灰色背包。红色、黑色、黑色、绿色和灰色这些颜色就是定性数据。

The data are the colors of houses. You sample five houses. The colors of the houses are white, yellow, white, red, and white. What type of data is this?

数据是房屋的颜色。你抽取了五所房屋。房屋颜色分别为白色、黄色、白色、红色和白色。这是什么类型的数据?

You may collect data as numbers and report it categorically. For example, the quiz scores for each student are recorded throughout the term. At the end of the term, the quiz scores are reported as A, B, C, D, or F.

你可以把数据作为数值收集,却以分类方式报告。例如,每个学生整个学期的测验分数都以数值记录。到学期结束时,测验分数被报告为 A、B、C、D 或 F。

Problem 问题

Work collaboratively to determine the correct data type (quantitative or qualitative). Indicate whether quantitative data are continuous or discrete. Hint: Data that are discrete often start with the words "the number of."

以合作方式确定正确的数据类型(定量或定性)。指出定量数据是连续的还是离散的。提示:离散数据常以"……的数量"开头。

1. the number of pairs of shoes you own

1. 你拥有的鞋的对数

2. the type of car you drive

2. 你所驾驶汽车的类型

3. the distance it is from your home to the nearest grocery store

3. 你家到最近的杂货店的距离

4. the number of classes you take per school year.

4. 你每个学年所修课程的数量。

5. the type of calculator you use

5. 你使用的计算器类型

6. weights of sumo wrestlers

6. 相扑选手的体重

7. number of correct answers on a quiz

7. 测验中答对题目的数量

8. IQ scores (This may cause some discussion.)

8. 智商分数(这一点可能会引发一些讨论。)

Solution 解答

Items a, d, and g are quantitative discrete; items c, f, and h are quantitative continuous; items b and e are qualitative, or categorical.

第 a、d、g 项为定量离散;第 c、f、h 项为定量连续;第 b、e 项为定性(或分类)数据。

Determine the correct data type (quantitative or qualitative) for the number of cars in a parking lot. Indicate whether quantitative data are continuous or discrete.

确定停车场中汽车数量的正确数据类型(定量或定性)。指出定量数据是连续的还是离散的。

Problem 问题

A statistics professor collects information about the classification of her students as freshmen, sophomores, juniors, or seniors. The data she collects are summarized in the pie chart Figure 1.3. What type of data does this graph show?

一位统计学教授收集了关于她的学生年级分类(大一、大二、大三或大四)的信息。她收集的数据在饼图 Figure 1.3 中进行了汇总。这张图显示的是什么类型的数据?

Solution 解答

This pie chart shows the students in each year, which is qualitative (or categorical) data.

这张饼图显示了每一年的学生人数,这属于定性(或分类)数据。

The registrar at State University keeps records of the number of credit hours students complete each semester. The data he collects are summarized in the histogram. The class boundaries are 10 to less than 13, 13 to less than 16, 16 to less than 19, 19 to less than 22, and 22 to less than 25.

州立大学的教务长保存着学生每个学期所修学分小时数的记录。他收集的数据在直方图中进行了汇总。组界为 10 至不足 13、13 至不足 16、16 至不足 19、19 至不足 22,以及 22 至不足 25。

What type of data does this graph show?

这张图显示的是什么类型的数据?

Qualitative Data Discussion 定性数据讨论

Below are tables comparing the number of part-time and full-time students at De Anza College and Foothill College enrolled for the spring 2010 quarter. The tables display counts (frequencies) and percentages or proportions (relative frequencies). The percent columns make comparing the same categories in the colleges easier. Displaying percentages along with the numbers is often helpful, but it is particularly important when comparing sets of data that do not have the same totals, such as the total enrollments for both colleges in this example. Notice how much larger the percentage for part-time students at Foothill College is compared to De Anza College.

下面这些表格比较了德安扎学院(De Anza College)和山麓学院(Foothill College)在 2010 年春季学期注册的兼职与全日制学生数量。这些表格显示了计数(频数)以及百分比或比例(相对频数)。百分比列使得比较两所学院中相同类别更为容易。同时显示百分比与数字通常是有帮助的,但在比较总数不同的数据集(如本例两所学院的总注册人数)时尤其重要。请注意,山麓学院兼职学生的百分比比德安扎学院大得多。

| De Anza College | | | Foothill College | | |

| 德安扎学院 | | | 山麓学院 | | |

|-----------------|--------|---------|------------------|--------|---------|

|-----------------|--------|---------|------------------|--------|---------|

| | Number | Percent | | Number | Percent |

| | 数量 | 百分比 | | 数量 | 百分比 |

| Full-time | 9,200 | 40.9% | Full-time | 4,059 | 28.6% |

| 全日制 | 9,200 | 40.9% | 全日制 | 4,059 | 28.6% |

| Part-time | 13,296 | 59.1% | Part-time | 10,124 | 71.4% |

| 兼职 | 13,296 | 59.1% | 兼职 | 10,124 | 71.4% |

| Total | 22,496 | 100% | Total | 14,183 | 100% |

| 总计 | 22,496 | 100% | 总计 | 14,183 | 100% |

Table 1.2 Fall Term 2007 (Census day)

表 1.2 2007 年秋季学期(普查日)

Tables are a good way of organizing and displaying data. But graphs can be even more helpful in understanding the data. There are no strict rules concerning which graphs to use. Two graphs that are used to display qualitative data are pie charts and bar graphs.

表格是组织和展示数据的一种好方法。但图形在理解数据方面可能更有帮助。关于使用哪些图形并没有严格的规定。用于展示定性数据的两种图形是饼图和条形图。

In a pie chart, categories of data are represented by wedges in a circle and are proportional in size to the percent of individuals in each category.

在饼图中,数据的各类别用圆中的扇形表示,其大小与各分类中个体所占百分比成正比。

In a bar graph, the length of the bar for each category is proportional to the number or percent of individuals in each category. Bars may be vertical or horizontal.

在条形图中,每一类别对应条形的长度与该类别中个体的数量或百分比成正比。条形可以是竖直的,也可以是水平的。

A Pareto chart consists of bars that are sorted into order by category size (largest to smallest).

帕累托图由按类别大小(从大到小)排序的条形组成。

Look at Figure 1.5 and Figure 1.6 and determine which graph (pie or bar) you think displays the comparisons better.

观察图 1.5 和图 1.6,判断你认为哪种图形(饼图或条形图)能更好地展示这些比较。

It is a good idea to look at a variety of graphs to see which is the most helpful in displaying the data. We might make different choices of what we think is the "best" graph depending on the data and the context. Our choice also depends on what we are using the data for.

最好查看多种图形,看看哪一种在展示数据方面最有帮助。根据数据和使用情境的不同,我们对于什么是"最佳"图形可能会做出不同的选择。我们的选择还取决于我们使用这些数据的目的。

Percentages That Add to More (or Less) Than 100% 百分比之和大于(或小于)100%

Sometimes percentages add up to be more than 100% (or less than 100%). In the graph, the percentages add to more than 100% because students can be in more than one category. A bar graph is appropriate to compare the relative size of the categories. A pie chart cannot be used. It also could not be used if the percentages added to less than 100%.

有时百分比之和会大于 100%(或小于 100%)。在该图中,百分比之和大于 100%,因为学生可能属于多个类别。条形图适合比较各类别的相对大小。饼图不能使用。如果百分比之和小于 100%,同样也不能使用饼图。

| Characteristic/Category | Percent |

| 特征/类别 | 百分比 |

|---------------------------------------------------------------------|---------|

|---------------------------------------------------------------------|---------|

| Full-Time Students | 40.9% |

| 全日制学生 | 40.9% |

| Students who intend to transfer to a 4-year educational institution | 48.6% |

| 打算转入四年制教育机构的学生 | 48.6% |

| Students under age 25 | 61.0% |

| 25 岁以下的学生 | 61.0% |

| TOTAL | 150.5% |

| 总计 | 150.5% |

Table 1.3 De Anza College Spring 2010

表 1.3 德安扎学院 2010 年春季

Omitting Categories/Missing Data 省略类别/缺失数据

The table displays Ethnicity of Students but is missing the "Other/Unknown" category. This category contains people who did not feel they fit into any of the ethnicity categories or declined to respond. Notice that the frequencies do not add up to the total number of students. In this situation, create a bar graph and not a pie chart.

该表显示了学生的种族,但缺失了"其他/未知"类别。这一类别包含了那些认为自己不属于任何种族类别或拒绝回答的人。注意,频数之和并不等于学生总数。在这种情况下,应绘制条形图而非饼图。

| | Frequency | Percent |

| | 频数 | 百分比 |

|------------------|----------------------|-------------------|

|------------------|----------------------|-------------------|

| Asian | 8,794 | 36.1% |

| 亚裔 | 8,794 | 36.1% |

| Black | 1,412 | 5.8% |

| 黑人 | 1,412 | 5.8% |

| Filipino | 1,298 | 5.3% |

| 菲律宾裔 | 1,298 | 5.3% |

| Hispanic | 4,180 | 17.1% |

| 西班牙裔 | 4,180 | 17.1% |

| Native American | 146 | 0.6% |

| 美洲原住民 | 146 | 0.6% |

| Pacific Islander | 236 | 1.0% |

| 太平洋岛民 | 236 | 1.0% |

| White | 5,978 | 24.5% |

| 白人 | 5,978 | 24.5% |

| TOTAL | 22,044 out of 24,382 | 90.4% out of 100% |

| 总计 | 22,044(共 24,382) | 90.4%(共 100%) |

Table 1.4 Ethnicity of Students at De Anza College Fall Term 2007 (Census Day)

表 1.4 德安扎学院学生种族 2007 年秋季学期(普查日)

The following graph is the same as the previous graph but the "Other/Unknown" percent (9.6%) has been included. The "Other/Unknown" category is large compared to some of the other categories (Native American, 0.6%, Pacific Islander 1.0%). This is important to know when we think about what the data are telling us.

下图与前图相同,但加入了"其他/未知"的百分比(9.6%)。与某些其他类别(美洲原住民 0.6%、太平洋岛民 1.0%)相比,"其他/未知"类别相当大。当我们思考这些数据在告诉我们什么时,这一点很重要。

This particular bar graph in Figure 1.9 can be difficult to understand visually. The graph in Figure 1.10 is a Pareto chart. The Pareto chart has the bars sorted from largest to smallest and is easier to read and interpret.

图 1.9 中这张特定的条形图在视觉上可能较难理解。图 1.10 中的图形是帕累托图。帕累托图的条形按从大到小排序,更易于阅读和解释。

Pie Charts: No Missing Data 饼图:无缺失数据

The following pie charts have the "Other/Unknown" category included (since the percentages must add to 100%). The chart in Figure 1.11(b) is organized by the size of each wedge, which makes it a more visually informative graph than the unsorted, alphabetical graph in Figure 1.11(a).

下面的饼图包含了"其他/未知"类别(因为百分比之和必须等于 100%)。图 1.11(b) 中的饼图按每个扇形的大小排列,因而比图 1.11(a) 中未按大小排序、按字母顺序排列的饼图在视觉上更具信息量。

Sampling 抽样

Gathering information about an entire population often costs too much or is virtually impossible. Instead, we use a sample of the population. A sample should have the same characteristics as the population it is representing. Most statisticians use various methods of random sampling in an attempt to achieve this goal. This section will describe a few of the most common methods. There are several different methods of random sampling. In each form of random sampling, each member of a population initially has an equal chance of being selected for the sample. Each method has pros and cons. The easiest method to describe is called a simple random sample. Any group of n individuals is equally likely to be chosen as any other group of n individuals if the simple random sampling technique is used. In other words, each sample of the same size has an equal chance of being selected. For example, suppose Lisa wants to form a four-person study group (herself and three other people) from her pre-calculus class, which has 31 members not including Lisa. To choose a simple random sample of size three from the other members of her class, Lisa could put all 31 names in a hat, shake the hat, close her eyes, and pick out three names. A more technological way is for Lisa to first list the last names of the members of her class together with a two-digit number, as in Table 1.5:

收集关于整个总体的信息往往成本过高,或者几乎不可能实现。相反,我们使用总体的一个样本。样本应当具有其所代表总体的相同特征。大多数统计学家使用各种随机抽样方法以试图达到这一目标。本节将描述几种最常用的方法。有几种不同的随机抽样方法。在每种随机抽样形式中,总体的每个成员最初都有相等的机会被选入样本。每种方法都有优缺点。最容易描述的方法称为简单随机样本。如果使用简单随机抽样技术,任意一组 n 个个体与任意另一组 n 个个体被选中的可能性相同。换言之,每一个同规模的样本都有相等的机会被选中。例如,假设 Lisa 想从她有 31 名成员(不含 Lisa 本人)的预修微积分课上组成一个四人学习小组(她自己和另外三人)。要从班上其他成员中抽取大小为三的简单随机样本,Lisa 可以把全部 31 个名字放进一顶帽子,摇晃帽子,闭上眼睛,抽出三个名字。一种更具技术性的方法是,Lisa 先把她班上成员的姓氏与一个两位数编号列在一起,如表 1.5 所示:
Table 1.5 Class Roster
IDNameIDNameIDName
00Anselmo11King21Roquero
01Bautista12Legeny22Roth
02Bayani13Lundquist23Rowell
03Cheng14Macierz24Salangsang
04Cuarismo15Motogawa25Slade
05Cuningham16Okimoto26Stratcher
06Fontecha17Patel27Tallai
07Hong18Price28Tran
08Hoobler19Quizon29Wai
09Jiao20Reyes30Wood
10Khan
表 1.5 班级名册
编号姓名编号姓名编号姓名
00Anselmo11King21Roquero
01Bautista12Legeny22Roth
02Bayani13Lundquist23Rowell
03Cheng14Macierz24Salangsang
04Cuarismo15Motogawa25Slade
05Cuningham16Okimoto26Stratcher
06Fontecha17Patel27Tallai
07Hong18Price28Tran
08Hoobler19Quizon29Wai
09Jiao20Reyes30Wood
10Khan

Lisa can use a table of random numbers (found in many statistics books and mathematical handbooks), a calculator, or a computer to generate random numbers. For this example, suppose Lisa chooses to generate random numbers from a calculator. The numbers generated are as follows:

Lisa 可以使用随机数表(见于许多统计学书籍和数学手册)、计算器或计算机来生成随机数。在本例中,假设 Lisa 选择用计算器生成随机数。生成的数列如下:

0.94360; 0.99832; 0.14669; 0.51470; 0.40581; 0.73381; 0.04399

0.94360;0.99832;0.14669;0.51470;0.40581;0.73381;0.04399

Lisa reads two-digit groups until she has chosen three class members (that is, she reads 0.94360 as the groups 94, 43, 36, 60). Each random number may only contribute one class member. If she needed to, Lisa could have generated more random numbers.

Lisa 每次读取两位数组,直到选满三名同学(即她把 0.94360 读作 94、43、36、60 这几组)。每个随机数最多只能贡献一名同学。如果需要,Lisa 本可以生成更多随机数。

The random numbers 0.94360 and 0.99832 do not contain appropriate two digit numbers. However the third random number, 0.14669, contains 14 (the fourth random number also contains 14), the fifth random number contains 05, and the seventh random number contains 04. The two-digit number 14 corresponds to Macierz, 05 corresponds to Cuningham, and 04 corresponds to Cuarismo. Besides herself, Lisa’s group will consist of Marcierz, Cuningham, and Cuarismo.

随机数 0.94360 和 0.99832 不含合适的两位数。然而第三个随机数 0.14669 含有 14(第四个随机数也含有 14),第五个随机数含有 05,第七个随机数含有 04。两位数 14 对应 Macierz,05 对应 Cuningham,04 对应 Cuarismo。除她自己外,Lisa 的小组将由 Marcierz、Cuningham 和 Cuarismo 组成。

To generate random numbers:

生成随机数的方法:

Note: randInt(0, 30, 3) will generate 3 random numbers.

注:randInt(0, 30, 3) 将生成 3 个随机数。

Besides simple random sampling, there are other forms of sampling that involve a chance process for getting the sample. Other well-known random sampling methods are the stratified sample, the cluster sample, and the systematic sample.

除简单随机抽样外,还有其他涉及机会过程来获取样本的抽样形式。其他著名的随机抽样方法有分层样本、整群样本和系统抽样。

To choose a stratified sample, divide the population into groups called strata and then take a proportionate number from each stratum. For example, you could stratify (group) your college population by department and then choose a proportionate simple random sample from each stratum (each department) to get a stratified random sample. To choose a simple random sample from each department, number each member of the first department, number each member of the second department, and do the same for the remaining departments. Then use simple random sampling to choose proportionate numbers from the first department and do the same for each of the remaining departments. Those numbers picked from the first department, picked from the second department, and so on represent the members who make up the stratified sample.

要抽取分层样本,先将总体划分为称为层的组,然后从每一层抽取成比例的数量。例如,你可以按院系对你的大学总体进行分层(分组),然后从每一层(每个院系)抽取成比例的一个简单随机样本,从而得到分层随机样本。要从每个院系抽取简单随机样本,先给第一个院系的每个成员编号,给第二个院系的每个成员编号,其余院系依此类推。然后使用简单随机抽样从第一个院系抽取成比例的数目,并对其余每个院系做同样的操作。从第一个院系、第二个院系……抽出的这些编号所代表的人即构成分层样本。

To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your college population, the four departments make up the cluster sample. Divide your college faculty by department. The departments are the clusters. Number each department, and then choose four different numbers using simple random sampling. All members of the four departments with those numbers are the cluster sample.

要抽取整群样本,先将总体划分为群(组),然后随机选中其中一些群。这些群中的所有成员都进入整群样本。例如,如果你从大学总体中随机抽取四个院系,这四个院系就构成整群样本。将大学教员按院系划分,院系即群。给每个院系编号,然后用简单随机抽样选出四个不同的编号。具有这些编号的四个院系中的所有成员即为整群样本。

To choose a systematic sample, randomly select a starting point and take every nth piece of data from a listing of the population. For example, suppose you have to do a phone survey. Your phone book contains 20,000 residence listings. You must choose 400 names for the sample. Number the population 1–20,000 and then use a simple random sample to pick a number that represents the first name in the sample. Then choose every fiftieth name thereafter until you have a total of 400 names (you might have to go back to the beginning of your phone list). Systematic sampling is frequently chosen because it is a simple method.

要抽取系统抽样,随机选择一个起点,然后从总体名单中每隔 n 个数据点取一个。例如,假设你要做一次电话调查。你的电话簿中有 20,000 条居民住址。你必须为样本选取 400 个名字。将总体编号为 1–20,000,然后用简单随机抽样选出一个代表样本中第一个名字的编号。之后每隔第五十个名字选取一个,直到总共得到 400 个名字(你可能需要回到电话名单的开头)。系统抽样之所以常被采用,是因为它是一种简单的方法。

A type of sampling that is non-random is convenience sampling. Convenience sampling involves using results that are readily available. For example, a computer software store conducts a marketing study by interviewing potential customers who happen to be in the store browsing through the available software. The results of convenience sampling may be very good in some cases and highly biased (favor certain outcomes) in others.

一种非随机的抽样类型是方便抽样。方便抽样使用随手可得的结果。例如,一家电脑软件商店通过对恰好在店里浏览可用软件的潜在顾客进行访谈来开展一项市场研究。方便抽样的结果在某些情况下可能很好,而在另一些情况下则可能高度偏倚(偏向某些结果)。

Sampling data should be done very carefully. Collecting data carelessly can have devastating results. Surveys mailed to households and then returned may be very biased (they may favor a certain group). It is better for the person conducting the survey to select the sample respondents.

抽样数据应当非常谨慎地进行。草率地收集数据可能导致灾难性的结果。邮寄到家庭并由其寄回的调查可能非常偏倚(它们可能偏向某一群体)。由实施调查的人来选择样本受访者更好。

True random sampling is done with replacement. That is, once a member is picked, that member goes back into the population and thus may be chosen more than once. However for practical reasons, in most populations, simple random sampling is done without replacement. Surveys are typically done without replacement. That is, a member of the population may be chosen only once. Most samples are taken from large populations and the sample tends to be small in comparison to the population. Since this is the case, sampling without replacement is approximately the same as sampling with replacement because the chance of picking the same individual more than once with replacement is very low.

真正的随机抽样是有放回进行的。也就是说,一旦某个成员被抽中,该成员就回到总体中,因此可能被选中不止一次。然而出于实际原因,在大多数总体中,简单随机抽样是无放回进行的。调查通常是无放回进行的。也就是说,总体中的某个成员只能被选中一次。大多数样本取自很大的总体,而样本相对于总体而言往往很小。既然如此,无放回抽样与有放回抽样近似相同,因为有放回抽样时同一人被选中不止一次的机会非常低。

Sampling without replacement instead of sampling with replacement becomes a mathematical issue only when the population is small.

只有当总体较小时,无放回抽样(相对于有放回抽样)才会成为一个数学问题。

When you analyze data, it is important to be aware of sampling errors and nonsampling errors. The actual process of sampling causes sampling errors. For example, the sample may not be large enough. Factors not related to the sampling process cause nonsampling errors. A defective counting device can cause a nonsampling error.

当你分析数据时,重要的是要注意抽样误差和非抽样误差。抽样这一实际过程会引起抽样误差。例如,样本可能不够大。与抽样过程无关的因素会引起非抽样误差。一个有缺陷的计数装置可能引起非抽样误差。

In reality, a sample will never be exactly representative of the population so there will always be some sampling error. As a rule, the larger the sample, the smaller the sampling error.

实际上,样本永远不会与总体完全代表一致,因此总会存在一些抽样误差。一般规律是,样本越大,抽样误差越小。

In statistics, a sampling bias is created when a sample is collected from a population and some members of the population are not as likely to be chosen as others (remember, each member of the population should have an equally likely chance of being chosen). When a sampling bias happens, there can be incorrect conclusions drawn about the population that is being studied.

在统计学中,当从总体中抽取样本而总体中的某些成员不像其他成员那样可能被选中时(记住,总体中的每个成员都应具有相等的被选中机会),就会产生抽样偏倚。当发生抽样偏倚时,就可能对所研究的总体得出不正确的结论。

Critical Evaluation 批判性评估

We need to evaluate the statistical studies we read about critically and analyze them before accepting the results of the studies. Common problems to be aware of include

我们需要批判性地评估并分析我们所读到的统计研究,在接受研究结果之前先加以分析。需要注意的常见问题包括:

As a class, determine whether or not the following samples are representative. If they are not, discuss the reasons.

作为一个班级,判断以下样本是否具有代表性。若不具有,讨论其原因。

1. To find the average GPA of all students in a university, use all honor students at the university as the sample.

1. 为求某大学全体学生的平均 GPA,以该大学所有荣誉学生作为样本。

2. To find out the most popular cereal among young people under the age of ten, stand outside a large supermarket for three hours and speak to every twentieth child under age ten who enters the supermarket.

2. 为查明十岁以下青少年中最受欢迎的麦片,在一家大型超市外站立三小时,与每一位进入超市的、未满十岁的第二十个孩子交谈。

3. To find the average annual income of all adults in the United States, sample U.S. congressmen. Create a cluster sample by considering each state as a stratum (group). By using simple random sampling, select states to be part of the cluster. Then survey every U.S. congressman in the cluster.

3. 为求美国所有成年人的平均年收入,对美国国会议员抽样。将每个州视为一个层(组)来构建整群样本。通过简单随机抽样,选取作为整群一部分的州。然后调查该整群中每一位美国国会议员。

4. To determine the proportion of people taking public transportation to work, survey 20 people in New York City. Conduct the survey by sitting in Central Park on a bench and interviewing every person who sits next to you.

4. 为确定乘坐公共交通上下班的人口比例,在纽约市调查 20 人。调查方式为坐在中央公园的一条长椅上,访谈坐在你旁边的每一个人。

5. To determine the average cost of a two-day stay in a hospital in Massachusetts, survey 100 hospitals across the state using simple random sampling.

5. 为确定马萨诸塞州医院两日住宿的平均费用,使用简单随机抽样对该州 100 家医院进行调查。

Problem 问题

A study is done to determine the average tuition that San Jose State undergraduate students pay per semester. Each student in the following samples is asked how much tuition he or she paid for the Fall semester. What is the type of sampling in each case?

一项研究旨在确定圣何塞州立大学本科生每学期支付的平均学费。询问以下样本中的每位学生其秋季学期支付了多少学费。每种情况下的抽样类型是什么?

1. A sample of 100 undergraduate San Jose State students is taken by organizing the students’ names by classification (freshman, sophomore, junior, or senior), and then selecting 25 students from each.

1. 将圣何塞州立大学学生的姓名按类别(大一、大二、大三或大四)整理,然后从每一类中选取 25 名学生,由此抽取 100 名圣何塞州立大学本科生的样本。

2. A random number generator is used to select a student from the alphabetical listing of all undergraduate students in the Fall semester. Starting with that student, every 50th student is chosen until 75 students are included in the sample.

2. 使用随机数生成器从秋季学期所有本科生的字母顺序名单中选出一个学生。从该学生开始,每隔第 50 个学生选取一个,直到样本包含 75 名学生。

3. A completely random method is used to select 75 students. Each undergraduate student in the fall semester has the same probability of being chosen at any stage of the sampling process.

3. 使用完全随机的方法选取 75 名学生。秋季学期的每位本科生在抽样过程的任何阶段都有相同的被选中概率。

4. The freshman, sophomore, junior, and senior years are numbered one, two, three, and four, respectively. A random number generator is used to pick two of those years. All students in those two years are in the sample.

4. 将大一、大二、大三、大四分别编号为 1、2、3、4。使用随机数生成器选出其中两年。这两年中的所有学生都进入样本。

5. An administrative assistant is asked to stand in front of the library one Wednesday and to ask the first 100 undergraduate students he encounters what they paid for tuition the Fall semester. Those 100 students are the sample.

5. 请一名行政助理在周三站在图书馆门前,询问他遇到的前 100 名本科生其秋季学期学费是多少。这 100 名学生即为样本。

Solution 解答

a\. stratified; b. systematic; c. simple random; d. cluster; e. convenience

a. 分层;b. 系统;c. 简单随机;d. 整群;e. 方便

You are going to use the random number generator to generate different types of samples from the data.

你将使用随机数生成器从这些数据中生成不同类型的样本。

This table displays six sets of quiz scores (each quiz counts 10 points) for an elementary statistics class.

该表显示了一个初等统计学班级的六组测验分数(每次测验满分 10 分)。
Table 1.6
#1#2#3#4#5#6
5710983
1059876
9108679
91010989
789574
9991087
7710988
8891088
978778
8810987
表 1.6
测验 1测验 2测验 3测验 4测验 5测验 6
5710983
1059876
9108679
91010989
789574
9991087
7710988
8891088
978778
8810987

Instructions: Use the Random Number Generator to pick samples.

说明:使用随机数生成器抽取样本。

1. Create a stratified sample by column. Pick three quiz scores randomly from each column.

1. 按列创建分层样本。从每一列随机抽取三个测验分数。

2. Create a cluster sample by picking two of the columns. Use the column numbers: one through six.

2. 通过选取其中两列来创建整群样本。使用列编号:1 到 6。

3. Create a simple random sample of 15 quiz scores.

3. 创建一个包含 15 个测验分数的简单随机样本。

4. Create a systematic sample of 12 quiz scores.

4. 创建一个包含 12 个测验分数的系统样本。

Problem 问题

Determine the type of sampling used (simple random, stratified, systematic, cluster, or convenience).

确定所使用的抽样类型(简单随机、分层、系统、整群或方便)。

1. A soccer coach selects six players from a group of boys aged eight to ten, seven players from a group of boys aged 11 to 12, and three players from a group of boys aged 13 to 14 to form a recreational soccer team.

1. 一名足球教练从一组 8 至 10 岁男孩中挑选 6 名球员,从一组 11 至 12 岁男孩中挑选 7 名,从一组 13 至 14 岁男孩中挑选 3 名,以组成一支休闲足球队。

2. A pollster interviews all human resource personnel in five different high tech companies.

2. 一名民意调查员访谈五家不同高科技公司中的所有人力资源人员。

3. A high school educational researcher interviews 50 high school female teachers and 50 high school male teachers.

3. 一名高中教育研究者访谈 50 名高中女教师和 50 名高中男教师。

4. A medical researcher interviews every third cancer patient from a list of cancer patients at a local hospital.

4. 一名医学研究者从当地一家医院癌症患者名单中每隔第三个癌症患者访谈一人。

5. A high school counselor uses a computer to generate 50 random numbers and then picks students whose names correspond to the numbers.

5. 一名高中辅导员用计算机生成 50 个随机数,然后挑选姓名对应这些数字的学生。

6. A student interviews classmates in his algebra class to determine how many pairs of jeans a student owns, on the average.

6. 一名学生访谈他代数课上的同学,以确定一名学生平均拥有多少条牛仔裤。

Solution 答案

a. stratified; b. cluster; c. stratified; d. systematic; e. simple random; f.convenience

a. 分层样本;b. 整群样本;c. 分层样本;d. 系统抽样;e. 简单随机样本;f. 方便样本

Determine the type of sampling used (simple random, stratified, systematic, cluster, or convenience).

确定所使用的抽样类型(简单随机、分层、系统、整群或方便)。

A high school principal polls 50 freshmen, 50 sophomores, 50 juniors, and 50 seniors regarding policy changes for after school activities.

一所高中的校长就课后活动政策变更,对 50 名高一、50 名高二、50 名高三和 50 名毕业班学生进行了调查。

If we were to examine two samples representing the same population, even if we used random sampling methods for the samples, they would not be exactly the same. Just as there is variation in data, there is variation in samples. As you become accustomed to sampling, the variability will begin to seem natural.

如果我们抽取两个代表同一总体的样本,即便使用了随机抽样方法,它们也不会完全相同。正如数据中存在变异,样本之间也存在变异。随着你逐渐习惯抽样,这种变异性会开始显得自然。

Suppose ABC College has 10,000 part-time students (the population). We are interested in the average amount of money a part-time student spends on books in the fall term. Asking all 10,000 students is an almost impossible task.

假设 ABC 学院有 10,000 名非全日制学生(总体)。我们关注一名非全日制学生在秋季学期购买书籍的平均花费金额。询问全部 10,000 名学生几乎是一项不可能完成的任务。

Suppose we take two different samples.

假设我们抽取两个不同的样本。

First, we use convenience sampling and survey ten students from a first term organic chemistry class. Many of these students are taking first term calculus in addition to the organic chemistry class. The amount of money they spend on books is as follows:

首先,我们使用方便抽样,调查了一个第一学期有机化学班中的十名学生。这些学生中许多人除了有机化学外还在修第一学期的微积分。他们在书籍上的花费如下:

\$128; \$87; \$173; \$116; \$130; \$204; \$147; \$189; \$93; \$153

128 美元;87 美元;173 美元;116 美元;130 美元;204 美元;147 美元;189 美元;93 美元;153 美元

The second sample is taken using a list of senior citizens who take P.E. classes and taking every fifth senior citizen on the list, for a total of ten senior citizens. They spend:

第二个样本取自一份参加体育课的老年人名单,在名单上每隔四位(即每第五位)老年人抽取一人,共十名老年人。他们的花费为:

\$50; \$40; \$36; \$15; \$50; \$100; \$40; \$53; \$22; \$22

50 美元;40 美元;36 美元;15 美元;50 美元;100 美元;40 美元;53 美元;22 美元;22 美元

It is unlikely that any student is in both samples.

任何一名学生同时出现在两个样本中的可能性很小。

Problem 问题

a. Do you think that either of these samples is representative of (or is characteristic of) the entire 10,000 part-time student population?

a. 你认为这两个样本中哪一个能代表(或具有)整个 10,000 名非全日制学生总体的特征?

Solution 答案

a. No. The first sample probably consists of science-oriented students. Besides the chemistry course, some of them are also taking first-term calculus. Books for these classes tend to be expensive. Most of these students are, more than likely, paying more than the average part-time student for their books. The second sample is a group of senior citizens who are, more than likely, taking courses for health and interest. The amount of money they spend on books is probably much less than the average parttime student. Both samples are biased. Also, in both cases, not all students have a chance to be in either sample.

a. 不能。第一个样本很可能由偏理科的学生组成。除了化学课,其中一些人还在修第一学期的微积分。这些课程的教材往往较贵。这些学生中的大多数很可能比非全日制学生的平均花费更多。第二个样本是一群老年人,他们很可能出于健康和兴趣而修课。他们在书籍上的花费很可能远低于非全日制学生的平均水平。两个样本都存在偏倚。而且,在两种情况下,并非所有学生都有机会进入任一样本。

Problem 问题

b. Since these samples are not representative of the entire population, is it wise to use the results to describe the entire population?

b. 既然这些样本不能代表整个总体,用这些结果来描述整个总体是否妥当?

Solution 答案

b. No. For these samples, each member of the population did not have an equally likely chance of being chosen.

b. 不妥。对于这些样本,总体中每个成员被抽中的机会并不均等。

Now, suppose we take a third sample. We choose ten different part-time students from the disciplines of chemistry, math, English, psychology, sociology, history, nursing, physical education, art, and early childhood development. (We assume that these are the only disciplines in which part-time students at ABC College are enrolled and that an equal number of part-time students are enrolled in each of the disciplines.) Each student is chosen using simple random sampling. Using a calculator, random numbers are generated and a student from a particular discipline is selected if he or she has a corresponding number. The students spend the following amounts:

现在,假设我们抽取第三个样本。我们从化学、数学、英语、心理学、社会学、历史、护理、体育、艺术和幼儿发展这些学科中各选取十名不同的非全日制学生。(我们假设这些是 ABC 学院非全日制学生注册的仅有的学科,并且每个学科注册的非全日制学生数量相等。)每名学生都通过简单随机抽样选出。使用计算器生成随机数,若某学科中的学生具有对应编号,则被选中。这些学生的花费如下:

\$180; \$50; \$150; \$85; \$260; \$75; \$180; \$200; \$200; \$150

180 美元;50 美元;150 美元;85 美元;260 美元;75 美元;180 美元;200 美元;200 美元;150 美元

Problem 问题

c. Is the sample biased?

c. 该样本是否存在偏倚?

Solution 答案

c. The sample is unbiased, but a larger sample would be recommended to increase the likelihood that the sample will be close to representative of the population. However, for a biased sampling technique, even a large sample runs the risk of not being representative of the population.

c. 该样本无偏倚,但建议增大样本量,以提高样本接近总体代表性的可能性。然而,对于存在偏倚的抽样技术,即使是大样本也存在不能代表总体的风险。

Students often ask if it is "good enough" to take a sample, instead of surveying the entire population. If the survey is done well, the answer is yes.

学生常会问,与其调查整个总体,抽取一个样本是否“足够好”。如果调查做得好,答案是肯定的。

A local radio station has a fan base of 20,000 listeners. The station wants to know if its audience would prefer more music or more talk shows. Asking all 20,000 listeners is an almost impossible task.

一家地方广播电台拥有 20,000 名听众的粉丝群。该电台想知道其受众是更偏好更多的音乐还是更多的谈话节目。询问全部 20,000 名听众几乎是一项不可能完成的任务。

The station uses convenience sampling and surveys the first 200 people they meet at one of the station’s music concert events. 24 people said they’d prefer more talk shows, and 176 people said they’d prefer more music.

该电台采用方便抽样,在电台的一场音乐演唱会活动中调查了他们遇到的前 200 人。24 人表示更偏好更多谈话节目,176 人表示更偏好更多音乐。

Do you think that this sample is representative of (or is characteristic of) the entire 20,000 listener population?

你认为这个样本能代表(或具有)整个 20,000 名听众总体的特征吗?

Variation in Data 数据中的变异

Variation is present in any set of data. For example, 16-ounce cans of beverage may contain more or less than 16 ounces of liquid. In one study, eight 16 ounce cans were measured and produced the following amount (in ounces) of beverage:

任何一组数据中都存在变异。例如,16 盎司的饮料罐中所含液体可能多于或少于 16 盎司。在一项研究中,测量了八罐 16 盎司饮料,得到如下饮料含量(单位:盎司):

15.8; 16.1; 15.2; 14.8; 15.8; 15.9; 16.0; 15.5

15.8;16.1;15.2;14.8;15.8;15.9;16.0;15.5

Measurements of the amount of beverage in a 16-ounce can may vary because different people make the measurements or because the exact amount, 16 ounces of liquid, was not put into the cans. Manufacturers regularly run tests to determine if the amount of beverage in a 16-ounce can falls within the desired range.

16 盎司饮料罐中饮料含量的测量值可能不同,原因可能是不同人员进行测量,也可能是由于罐中实际灌入的液体并非精确的 16 盎司。制造商会定期进行测试,以确定 16 盎司饮料罐中的饮料含量是否落在期望范围内。

Be aware that as you take data, your data may vary somewhat from the data someone else is taking for the same purpose. This is completely natural. However, if two or more of you are taking the same data and get very different results, it is time for you and the others to reevaluate your data-taking methods and your accuracy.

请注意,在你收集数据时,你的数据可能与他人出于相同目的所收集的数据略有不同。这完全是自然的。然而,如果你们两人或多人收集同一数据时得到了差异很大的结果,那么你和他人就应当重新评估你们的收集方法和准确性。

Variation in Samples 样本中的变异

It was mentioned previously that two or more samples from the same population, taken randomly, and having close to the same characteristics of the population will likely be different from each other. Suppose Doreen and Jung both decide to study the average amount of time students at their college sleep each night. Doreen and Jung each take samples of 500 students. Doreen uses systematic sampling and Jung uses cluster sampling. Doreen's sample will be different from Jung's sample. Even if Doreen and Jung used the same sampling method, in all likelihood their samples would be different. Neither would be wrong, however.

前面提到过,从同一总体中随机抽取、且特征与总体相近的两个或多个样本,彼此之间很可能不同。假设 Doreen 和 Jung 都决定研究他们学院学生每晚的平均睡眠时间。Doreen 和 Jung 各自抽取了 500 名学生的样本。Doreen 使用系统抽样,Jung 使用整群抽样。Doreen 的样本将与 Jung 的样本不同。即便 Doreen 和 Jung 使用相同的抽样方法,他们的样本也很可能不同。然而,两者都不会是错的。

Think about what contributes to making Doreen’s and Jung’s samples different.

思考一下,是什么导致了 Doreen 和 Jung 的样本存在差异。

If Doreen and Jung took larger samples (i.e. the number of data values is increased), their sample results (the average amount of time a student sleeps) might be closer to the actual population average. But still, their samples would be, in all likelihood, different from each other. This variability in samples cannot be stressed enough.

如果 Doreen 和 Jung 抽取更大的样本(即数据值的数量增加),他们的样本结果(一名学生的平均睡眠时间)可能会更接近真实的总体均值。但即便如此,他们的样本在很大程度上仍会彼此不同。这种样本变异性再怎么强调也不为过。

Size of a Sample 样本量

The size of a sample (often called the number of observations) is important. The examples you have seen in this book so far have been small. Samples of only a few hundred observations, or even smaller, are sufficient for many purposes. In polling, samples that are from 1,200 to 1,500 observations are considered large enough and good enough if the survey is random and is well done. You will learn why when you study confidence intervals.

样本量(常称为观测数)很重要。你在本书中迄今看到的例子规模都很小。仅有几百个观测值甚至更小的样本,对许多目的来说已经足够。在民意调查中,如果调查是随机且执行良好的,1,200 到 1,500 个观测值的样本就被认为是足够大且足够好的。当你学习置信区间时就会明白原因。

Be aware that many large samples are biased. For example, call-in surveys are invariably biased, because people choose to respond or not.

请注意,许多大样本存在偏倚。例如,来电调查总是存在偏倚,因为人们可以选择是否回应。

Divide into groups of two, three, or four. Your instructor will give each group one six-sided die. Try this experiment twice. Roll one fair die (six-sided) 20 times. Record the number of ones, twos, threes, fours, fives, and sixes you get in Table 1.7 and Table 1.8 (“frequency” is the number of times a particular face of the die occurs):

分成二、三或四人一组。你的教师会给每组一枚六面骰子。把这个实验做两遍。将一枚均匀的骰子(六面)投掷 20 次。把你在表 1.7 和表 1.8 中得到的 1 点、2 点、3 点、4 点、5 点和 6 点的次数记录下来(“频数”指骰子某一面出现的次数):
Table 1.7 First Experiment (20 rolls)
Face on DieFrequency
1
2
3
4
5
6
表 1.7 第一次实验(20 次投掷)
骰子面频数
1
2
3
4
5
6
Table 1.8 Second Experiment (20 rolls)
Face on DieFrequency
1
2
3
4
5
6
表 1.8 第二次实验(20 次投掷)
骰子面频数
1
2
3
4
5
6

Did the two experiments have the same results? Probably not. If you did the experiment a third time, do you expect the results to be identical to the first or second experiment? Why or why not?

两次实验的结果相同吗?大概不会。如果你再做第三次实验,你预期结果会与第一次或第二次完全相同吗?为什么或为什么不?

Which experiment had the correct results? They both did. The job of the statistician is to see through the variability and draw appropriate conclusions.

哪一次实验的结果是正确的?两次都是正确的。统计学家的职责是看透变异性并得出恰当的结论。

1.3 Frequency, Frequency Tables, and Levels of Measurement 1.3 频数、频数表与测量尺度

Once you have a set of data, you will need to organize it so that you can analyze how frequently each datum occurs in the set. However, when calculating the frequency, you may need to round your answers so that they are as precise as possible.

一旦得到一组数据,就需要对其进行整理,以便分析每个数据点在数据集中出现的频繁程度。然而,在计算频数时,你可能需要对结果进行四舍五入,使其尽可能精确。

Answers and Rounding Off 答案与四舍五入

A simple way to round off answers is to carry your final answer one more decimal place than was present in the original data. Round off only the final answer. Do not round off any intermediate results, if possible. If it becomes necessary to round off intermediate results, carry them to at least twice as many decimal places as the final answer. For example, the average of the three quiz scores four, six, and nine is 6.3, rounded off to the nearest tenth, because the data are whole numbers. Most answers will be rounded off in this manner.

一种简单的四舍五入方法是:最终结果比原始数据多保留一位小数。只对最终结果进行四舍五入。如果可能,不要对任何中间结果进行四舍五入。如果必须对中间结果四舍五入,则应至少保留到最终结果小数位数的两倍。例如,三次测验成绩 4、6、9 的均值为 6.3,四舍五入到十分位,因为数据为整数。大多数答案都将以这种方式四舍五入。

It is not necessary to reduce most fractions in this course. Especially in Probability Topics, the chapter on probability, it is more helpful to leave an answer as an unreduced fraction.

在本课程中,大多数分数无需约分。尤其是在“概率主题”这一概率章中,将答案保留为未约分的分数更有帮助。

Levels of Measurement 测量尺度

The way a set of data is measured is called its level of measurement. Correct statistical procedures depend on a researcher being familiar with levels of measurement. Not every statistical operation can be used with every set of data. Data can be classified into four levels of measurement. They are (from lowest to highest level):

对一组数据进行测量的方式称为其测量尺度。正确的统计方法依赖于研究者熟悉测量尺度。并非每一种统计操作都适用于每一组数据。数据可分为四个测量尺度,它们(从最低到最高)是:

Data that is measured using a nominal scale is qualitative (categorical). Categories, colors, names, labels and favorite foods along with yes or no responses are examples of nominal level data. Nominal scale data are not ordered. For example, trying to classify people according to their favorite food does not make any sense. Putting pizza first and sushi second is not meaningful.

用名义尺度测量的数据是定性(分类)数据。类别、颜色、名称、标签以及最喜欢的食物,连同“是”或“否”的回答,都是名义尺度数据的例子。名义尺度数据没有顺序。例如,试图按人们最喜欢的食物对他们分类毫无意义。把披萨排第一、寿司排第二是没有意义的。

Smartphone companies are another example of nominal scale data. The data are the names of the companies that make smartphones, but there is no agreed upon order of these brands, even though people may have personal preferences. Nominal scale data cannot be used in calculations.

智能手机制造商是名义尺度数据的另一个例子。数据是制造智能手机的公司名称,但这些品牌之间没有公认的排序,尽管人们可能有个人偏好。名义尺度数据不能用于计算。

Data that is measured using an ordinal scale is similar to nominal scale data but there is a big difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks in the United States. The top five national parks in the United States can be ranked from one to five but we cannot measure differences between the data.

用顺序尺度测量的数据与名义尺度数据相似,但有一个重大区别。顺序尺度数据可以排序。顺序尺度数据的一个例子是美国排名前五的国家公园列表。美国排名前五的国家公园可以从第一排到第五,但我们无法测量数据之间的差异。

Another example of using the ordinal scale is a cruise survey where the responses to questions about the cruise are “excellent,” “good,” “satisfactory,” and “unsatisfactory.” These responses are ordered from the most desired response to the least desired. But the differences between two pieces of data cannot be measured. Like the nominal scale data, ordinal scale data cannot be used in calculations.

使用顺序尺度的另一个例子是一项邮轮调查,其中关于邮轮的问卷回答为“极好”“好”“满意”和“不满意”。这些回答从最受期望到最不受期望进行了排序。但两段数据之间的差异无法测量。与名义尺度数据一样,顺序尺度数据也不能用于计算。

Data that is measured using the interval scale is similar to ordinal level data because it has a definite ordering but there is a difference between data. The differences between interval scale data can be measured though the data does not have a starting point.

用区间尺度测量的数据与顺序尺度数据相似,因为它有明确的排序,但数据之间存在差异。区间尺度数据之间的差异可以测量,尽管该数据没有起点。

Temperature scales like Celsius (C) and Fahrenheit (F) are measured by using the interval scale. In both temperature measurements, 40° is equal to 100° minus 60°. Differences make sense. But 0 degrees does not because, in both scales, 0 is not the absolute lowest temperature. Temperatures like -10° F and -15° C exist and are colder than 0.

摄氏温标(C)和华氏温标(F)等温度刻度使用区间尺度测量。在两种温度测量中,40° 等于 100° 减去 60°。差值是有意义的。但 0 度没有意义,因为在这两种温标中,0 都不是绝对最低温度。像 -10° F 和 -15° C 这样的温度是存在的,并且比 0 更冷。

Interval level data can be used in calculations, but one type of comparison cannot be done. 80° C is not four times as hot as 20° C (nor is 80° F four times as hot as 20° F). There is no meaning to the ratio of 80 to 20 (or four to one).

区间尺度数据可用于计算,但有一种比较无法完成。80° C 的热度不是 20° C 的四倍(80° F 的热度也不是 20° F 的四倍)。80 与 20(或四与一)的比值没有意义。

Data that is measured using the ratio scale takes care of the ratio problem and gives you the most information. Ratio scale data is like interval scale data, but it has a 0 point and ratios can be calculated. For example, four multiple choice statistics final exam scores are 80, 68, 20 and 92 (out of a possible 100 points). The exams are machine-graded.

用比率尺度测量的数据解决了比值问题,并为你提供最多的信息。比率尺度数据类似于区间尺度数据,但它有一个 0 点,且可以计算比值。例如,四次统计学选择题期末考试成绩分别为 80、68、20 和 92(满分 100 分)。这些考试由机器阅卷。

The data can be put in order from lowest to highest: 20, 68, 80, 92.

数据可以按从低到高排序:20、68、80、92。

The differences between the data have meaning. The score 92 is more than the score 68 by 24 points. Ratios can be calculated. The smallest score is 0. So 80 is four times 20. The score of 80 is four times better than the score of 20.

数据之间的差异是有意义的。92 分比 68 分高出 24 分。可以计算比值。最低分数为 0。因此 80 是 20 的四倍。80 分的成绩是 20 分的四倍。

Frequency 频数

Twenty students were asked how many hours they worked per day. Their responses, in hours, are as follows: 5; 6; 3; 3; 2; 4; 7; 5; 2; 3; 5; 6; 5; 4; 4; 3; 5; 2; 5; 3.

有人询问 20 名学生每天工作多少小时。他们的回答(以小时计)如下:5; 6; 3; 3; 2; 4; 7; 5; 2; 3; 5; 6; 5; 4; 4; 3; 5; 2; 5; 3。

Table 1.9 lists the different data values in ascending order and their frequencies.

表 1.9 按升序列出各个不同的数据值及其频数。
DATA VALUE FREQUENCY
2 3
3 5
4 3
5 6
6 2
7 1
数据值 频数
2 3
3 5
4 3
5 6
6 2
7 1

Table 1.9 Frequency Table of Student Work Hours

表 1.9 学生工作时长频数表

A frequency is the number of times a value of the data occurs. According to Table 1.9, there are three students who work two hours, five students who work three hours, and so on. The sum of the values in the frequency column, 20, represents the total number of students included in the sample.

频数是某个数据值出现的次数。根据表 1.9,有 3 名学生工作 2 小时,5 名学生工作 3 小时,依此类推。频数列各值之和为 20,即样本中包含的学生总数。

A relative frequency is the ratio (fraction or proportion) of the number of times a value of the data occurs in the set of all outcomes to the total number of outcomes. To find the relative frequencies, divide each frequency by the total number of students in the sample–in this case, 20. Relative frequencies can be written as fractions, percents, or decimals.

相对频数是某个数据值在全部结果中出现的次数与结果总数之比(分数或比例)。求相对频数时,把每个频数除以样本中的学生总数——本例中为 20。相对频数可写成分数、百分比或小数。
DATA VALUE FREQUENCY RELATIVE FREQUENCY
2 3 $\frac{3}{20}$ or 0.15
3 5 $\frac{5}{20}$ or 0.25
4 3 $\frac{3}{20}$ or 0.15
5 6 $\frac{6}{20}$ or 0.30
6 2 $\frac{2}{20}$ or 0.10
7 1 $\frac{1}{20}$ or 0.05
数据值 频数 相对频数
2 3 $\frac{3}{20}$ 或 0.15
3 5 $\frac{5}{20}$ 或 0.25
4 3 $\frac{3}{20}$ 或 0.15
5 6 $\frac{6}{20}$ 或 0.30
6 2 $\frac{2}{20}$ 或 0.10
7 1 $\frac{1}{20}$ 或 0.05

Table 1.10 Frequency Table of Student Work Hours with Relative Frequencies

表 1.10 含相对频数的学生工作时长频数表

The sum of the values in the relative frequency column of Table 1.10 is $\frac{20}{20}$ , or 1.

表 1.10 相对频数列各值之和为 $\frac{20}{20}$ ,即 1。

Cumulative relative frequency is the accumulation of the previous relative frequencies. To find the cumulative relative frequencies, add all the previous relative frequencies to the relative frequency for the current row, as shown in Table 1.11.

累积相对频数是此前各相对频数的累加。求累积相对频数时,把此前所有相对频数与当前行的相对频数相加,如表 1.11 所示。
DATA VALUE FREQUENCY RELATIVE
FREQUENCY
CUMULATIVE RELATIVE
FREQUENCY
2 3 $\frac{3}{20}$ or 0.15 0.15
3 5 $\frac{5}{20}$ or 0.25 0.15 + 0.25 = 0.40
4 3 $\frac{3}{20}$ or 0.15 0.40 + 0.15 = 0.55
5 6 $\frac{6}{20}$ or 0.30 0.55 + 0.30 = 0.85
6 2 $\frac{2}{20}$ or 0.10 0.85 + 0.10 = 0.95
7 1 $\frac{1}{20}$ or 0.05 0.95 + 0.05 = 1.00
数据值 频数 相对
频数
累积相对
频数
2 3 $\frac{3}{20}$ 或 0.15 0.15
3 5 $\frac{5}{20}$ 或 0.25 0.15 + 0.25 = 0.40
4 3 $\frac{3}{20}$ 或 0.15 0.40 + 0.15 = 0.55
5 6 $\frac{6}{20}$ 或 0.30 0.55 + 0.30 = 0.85
6 2 $\frac{2}{20}$ 或 0.10 0.85 + 0.10 = 0.95
7 1 $\frac{1}{20}$ 或 0.05 0.95 + 0.05 = 1.00

Table 1.11 Frequency Table of Student Work Hours with Relative and Cumulative Relative Frequencies

表 1.11 含相对频数与累积相对频数的学生工作时长频数表

The last entry of the cumulative relative frequency column is one, indicating that one hundred percent of the data has been accumulated.

累积相对频数列的最后一项为 1,表示已累计全部数据的百分之百。

Because of rounding, the relative frequency column may not always sum to one, and the last entry in the cumulative relative frequency column may not be one. However, they each should be close to one.

由于四舍五入,相对频数列之和未必总是 1,累积相对频数列的最后一项也未必是 1,但二者都应接近 1。

Table 1.12 represents the heights, in inches, of a sample of 100 male semiprofessional soccer players.

表 1.12 给出 100 名男性半职业足球运动员样本的身高(以英寸计)。
HEIGHTS
(INCHES)
FREQUENCY RELATIVE
FREQUENCY
CUMULATIVE
RELATIVE
FREQUENCY
59.95–61.95 5 $\frac{5}{100}$ = 0.05 0.05
61.95–63.95 3 $\frac{3}{100}$ = 0.03 0.05 + 0.03 = 0.08
63.95–65.95 15 $\frac{15}{100}$ = 0.15 0.08 + 0.15 = 0.23
65.95–67.95 40 $\frac{40}{100}$ = 0.40 0.23 + 0.40 = 0.63
67.95–69.95 17 $\frac{17}{100}$ = 0.17 0.63 + 0.17 = 0.80
69.95–71.95 12 $\frac{12}{100}$ = 0.12 0.80 + 0.12 = 0.92
71.95–73.95 7 $\frac{7}{100}$ = 0.07 0.92 + 0.07 = 0.99
73.95–75.95 1 $\frac{1}{100}$ = 0.01 0.99 + 0.01 = 1.00
Total = 100 Total = 1.00
身高
(英寸)
频数 相对
频数
累积
相对
频数
59.95–61.95 5 $\frac{5}{100}$ = 0.05 0.05
61.95–63.95 3 $\frac{3}{100}$ = 0.03 0.05 + 0.03 = 0.08
63.95–65.95 15 $\frac{15}{100}$ = 0.15 0.08 + 0.15 = 0.23
65.95–67.95 40 $\frac{40}{100}$ = 0.40 0.23 + 0.40 = 0.63
67.95–69.95 17 $\frac{17}{100}$ = 0.17 0.63 + 0.17 = 0.80
69.95–71.95 12 $\frac{12}{100}$ = 0.12 0.80 + 0.12 = 0.92
71.95–73.95 7 $\frac{7}{100}$ = 0.07 0.92 + 0.07 = 0.99
73.95–75.95 1 $\frac{1}{100}$ = 0.01 0.99 + 0.01 = 1.00
合计 = 100 合计 = 1.00

Table 1.12 Frequency Table of Soccer Player Height

表 1.12 足球运动员身高频数表

The data in this table have been grouped into the following intervals:

该表中的数据已分组为以下区间:

This example is used again in Descriptive Statistics, where the method used to compute the intervals will be explained.

本例将在“描述统计学”一章中再次使用,届时会说明区间的计算方法。

In this sample, there are five players whose heights fall within the interval 59.95–61.95 inches, three players whose heights fall within the interval 61.95–63.95 inches, 15 players whose heights fall within the interval 63.95–65.95 inches, 40 players whose heights fall within the interval 65.95–67.95 inches, 17 players whose heights fall within the interval 67.95–69.95 inches, 12 players whose heights fall within the interval 69.95–71.95, seven players whose heights fall within the interval 71.95–73.95, and one player whose heights fall within the interval 73.95–75.95. All heights fall between the endpoints of an interval and not at the endpoints.

在这个样本中,有 5 名运动员的身高落在 59.95–61.95 英寸区间内,3 名落在 61.95–63.95 英寸区间内,15 名落在 63.95–65.95 英寸区间内,40 名落在 65.95–67.95 英寸区间内,17 名落在 67.95–69.95 英寸区间内,12 名落在 69.95–71.95 区间内,7 名落在 71.95–73.95 区间内,1 名落在 73.95–75.95 区间内。所有身高都落在区间内部,而不在区间端点上。

Problem 问题

From Table 1.12, find the percentage of heights that are less than 65.95 inches.

根据表 1.12,求身高小于 65.95 英寸的百分比。

Solution 解答

If you look at the first, second, and third rows, the heights are all less than 65.95 inches. There are 5 + 3 + 15 = 23 players whose heights are less than 65.95 inches. The percentage of heights less than 65.95 inches is then $\frac{23}{100}$ or 23%. This percentage is the cumulative relative frequency entry in the third row.

看第一、第二和第三行,这些身高都小于 65.95 英寸。身高小于 65.95 英寸的运动员共有 5 + 3 + 15 = 23 名。于是身高小于 65.95 英寸的百分比为 $\frac{23}{100}$ 或 23%。这个百分比正是第三行的累积相对频数。

Table 1.13 shows the amount, in inches, of annual rainfall in a sample of towns.

表 1.13 给出一个城镇样本的年降雨量(以英寸计)。
Rainfall (Inches) Frequency Relative Frequency Cumulative Relative Frequency
2.95–4.97 6 $\frac{6}{50}$ = 0.12 0.12
4.97–6.99 7 $\frac{7}{50}$ = 0.14 0.12 + 0.14 = 0.26
6.99–9.01 15 $\frac{15}{50}$ = 0.30 0.26 + 0.30 = 0.56
9.01–11.03 8 $\frac{8}{50}$ = 0.16 0.56 + 0.16 = 0.72
11.03–13.05 9 $\frac{9}{50}$ = 0.18 0.72 + 0.18 = 0.90
13.05–15.07 5 $\frac{5}{50}$ = 0.10 0.90 + 0.10 = 1.00
Total = 50 Total = 1.00
降雨量(英寸) 频数 相对频数 累积相对频数
2.95–4.97 6 $\frac{6}{50}$ = 0.12 0.12
4.97–6.99 7 $\frac{7}{50}$ = 0.14 0.12 + 0.14 = 0.26
6.99–9.01 15 $\frac{15}{50}$ = 0.30 0.26 + 0.30 = 0.56
9.01–11.03 8 $\frac{8}{50}$ = 0.16 0.56 + 0.16 = 0.72
11.03–13.05 9 $\frac{9}{50}$ = 0.18 0.72 + 0.18 = 0.90
13.05–15.07 5 $\frac{5}{50}$ = 0.10 0.90 + 0.10 = 1.00
合计 = 50 合计 = 1.00

Table 1.13

表 1.13

From Table 1.13, find the percentage of rainfall that is less than 9.01 inches.

根据表 1.13,求降雨量小于 9.01 英寸的百分比。

Problem 问题

From Table 1.12, find the percentage of heights that fall between 61.95 and 65.95 inches.

根据表 1.12,求身高落在 61.95 与 65.95 英寸之间的百分比。

Solution 解答

Add the relative frequencies in the second and third rows: 0.03 + 0.15 = 0.18 or 18%.

把第二行与第三行的相对频数相加:0.03 + 0.15 = 0.18,即 18%。

From Table 1.13, find the percentage of rainfall that is between 6.99 and 13.05 inches.

根据表 1.13,求降雨量在 6.99 与 13.05 英寸之间的百分比。

Problem 问题

Use the heights of the 100 male semiprofessional soccer players in Table 1.12. Fill in the blanks and check your answers.

使用表 1.12 中 100 名男性半职业足球运动员的身高。填空并检验你的答案。

1. The percentage of heights that are from 67.95 to 71.95 inches is: \_\_\_\_.

1. 身高在 67.95 至 71.95 英寸之间的百分比是:\_\_\_\_。

2. The percentage of heights that are from 67.95 to 73.95 inches is: \_\_\_\_.

2. 身高在 67.95 至 73.95 英寸之间的百分比是:\_\_\_\_。

3. The percentage of heights that are more than 65.95 inches is: \_\_\_\_.

3. 身高大于 65.95 英寸的百分比是:\_\_\_\_。

4. The number of players in the sample who are between 61.95 and 71.95 inches tall is: \_\_\_\_.

4. 样本中身高在 61.95 与 71.95 英寸之间的运动员人数是:\_\_\_\_。

5. What kind of data are the heights?

5. 身高属于哪种数据?

6. Describe how you could gather this data (the heights) so that the data are characteristic of all male semiprofessional soccer players.

6. 说明你会如何收集这些数据(身高),使数据能代表所有男性半职业足球运动员。

Remember, you count frequencies. To find the relative frequency, divide the frequency by the total number of data values. To find the cumulative relative frequency, add all of the previous relative frequencies to the relative frequency for the current row.

记住,频数是数出来的。求相对频数时,把频数除以数据值总数。求累积相对频数时,把此前所有相对频数与当前行的相对频数相加。

Solution 解答

1. 29%

1. 29%

2. 36%

2. 36%

3. 77%

3. 77%

4. 87

4. 87

5. quantitative continuous

5. 定量连续数据

6. get rosters from each team and choose a simple random sample from each

6. 从每支球队取得球员名单,并从每份名单中抽取简单随机样本

From Table 1.13, find the number of towns that have rainfall between 2.95 and 9.01 inches.

根据表 1.13,求降雨量在 2.95 与 9.01 英寸之间的城镇数目。

In your class, have someone conduct a survey of the number of siblings (brothers and sisters) each student has. Create a frequency table. Add to it a relative frequency column and a cumulative relative frequency column. Answer the following questions:

在你的班级里,请一位同学调查每名学生有多少兄弟姐妹。制作一张频数表,并为它添加相对频数列和累积相对频数列。回答以下问题:

1. What percentage of the students in your class have no siblings?

1. 你班上有多少百分比的学生没有兄弟姐妹?

2. What percentage of the students have from one to three siblings?

2. 有多少百分比的学生有 1 至 3 个兄弟姐妹?

3. What percentage of the students have fewer than three siblings?

3. 有多少百分比的学生的兄弟姐妹少于 3 个?

Nineteen people were asked how many miles, to the nearest mile, they commute to work each day. The data are as follows: 2; 5; 7; 3; 2; 10; 18; 15; 20; 7; 10; 18; 5; 12; 13; 12; 4; 5; 10. Table 1.14 was produced:

有人询问 19 个人每天通勤上班的路程为多少英里(取到最近的整英里)。数据如下:2; 5; 7; 3; 2; 10; 18; 15; 20; 7; 10; 18; 5; 12; 13; 12; 4; 5; 10。据此得到表 1.14:
DATA FREQUENCY RELATIVE
FREQUENCY
CUMULATIVE
RELATIVE
FREQUENCY
3 3 $\frac{3}{19}$ 0.1579
4 1 $\frac{1}{19}$ 0.2105
5 3 $\frac{3}{19}$ 0.1579
7 2 $\frac{2}{19}$ 0.2632
10 3 $\frac{4}{19}$ 0.4737
12 2 $\frac{2}{19}$ 0.7895
13 1 $\frac{1}{19}$ 0.8421
15 1 $\frac{1}{19}$ 0.8948
18 1 $\frac{1}{19}$ 0.9474
20 1 $\frac{1}{19}$ 1.0000
数据 频数 相对
频数
累积
相对
频数
3 3 $\frac{3}{19}$ 0.1579
4 1 $\frac{1}{19}$ 0.2105
5 3 $\frac{3}{19}$ 0.1579
7 2 $\frac{2}{19}$ 0.2632
10 3 $\frac{4}{19}$ 0.4737
12 2 $\frac{2}{19}$ 0.7895
13 1 $\frac{1}{19}$ 0.8421
15 1 $\frac{1}{19}$ 0.8948
18 1 $\frac{1}{19}$ 0.9474
20 1 $\frac{1}{19}$ 1.0000

Table 1.14 Frequency of Commuting Distances

表 1.14 通勤距离频数表

Problem 问题

1. Is the table correct? If it is not correct, what is wrong?

1. 这张表正确吗?如果不正确,错在哪里?

2. True or False: Three percent of the people surveyed commute three miles. If the statement is not correct, what should it be? If the table is incorrect, make the corrections.

2. 判断对错:受访者中有百分之三的人通勤 3 英里。如果这句话不对,正确的说法应当是什么?如果表格有误,请予以更正。

3. What fraction of the people surveyed commute five or seven miles?

3. 受访者中通勤 5 英里或 7 英里的人占的分数是多少?

4. What fraction of the people surveyed commute 12 miles or more? Less than 12 miles? Between five and 13 miles (not including five and 13 miles)?

4. 受访者中通勤 12 英里及以上的人占的分数是多少?少于 12 英里的呢?在 5 英里与 13 英里之间(不含 5 英里和 13 英里)的呢?

Solution 解答

1. No. The frequency column sums to 18, not 19. Not all cumulative relative frequencies are correct.

1. 不正确。频数列之和为 18,而不是 19。并非所有累积相对频数都正确。

2. False. The frequency for three miles should be one; for two miles (left out), two. The cumulative relative frequency column should read: 0.1052, 0.1579, 0.2105, 0.3684, 0.4737, 0.6316, 0.7368, 0.7895, 0.8421, 0.9474, 1.0000.

2. 错。3 英里的频数应为 1;被遗漏的 2 英里的频数应为 2。累积相对频数列应为:0.1052, 0.1579, 0.2105, 0.3684, 0.4737, 0.6316, 0.7368, 0.7895, 0.8421, 0.9474, 1.0000。

3. $\frac{5}{19}$

3. $\frac{5}{19}$

4. $\frac{7}{19}$, $\frac{12}{19}$, $\frac{7}{19}$

4. $\frac{7}{19}$,$\frac{12}{19}$,$\frac{7}{19}$

Table 1.13 represents the amount, in inches, of annual rainfall in a sample of towns. What fraction of towns surveyed get between 11.03 and 13.05 inches of rainfall each year?

表 1.13 给出一个城镇样本的年降雨量(以英寸计)。所调查的城镇中,每年降雨量在 11.03 与 13.05 英寸之间的占的分数是多少?

Table 1.15 contains the total number of deaths worldwide as a result of earthquakes for the period from 2000 to 2012.

表 1.15 列出 2000 年至 2012 年期间全世界因地震死亡的总人数。
Year Total Number of Deaths
2000 231
2001 21,357
2002 11,685
2003 33,819
2004 228,802
2005 88,003
2006 6,605
2007 712
2008 88,011
2009 1,790
2010 320,120
2011 21,953
2012 768
Total 823,856
年份 死亡总人数
2000 231
2001 21,357
2002 11,685
2003 33,819
2004 228,802
2005 88,003
2006 6,605
2007 712
2008 88,011
2009 1,790
2010 320,120
2011 21,953
2012 768
合计 823,856

Table 1.15

表 1.15

Problem 问题

Answer the following questions.

回答以下问题。

1. What is the frequency of deaths measured from 2006 through 2009?

1. 2006 年至 2009 年期间统计到的死亡频数是多少?

2. What percentage of deaths occurred after 2009?

2. 2009 年之后发生的死亡占多少百分比?

3. What is the relative frequency of deaths that occurred in 2003 or earlier?

3. 2003 年及以前发生的死亡的相对频数是多少?

4. What is the percentage of deaths that occurred in 2004?

4. 2004 年发生的死亡占多少百分比?

5. What kind of data are the numbers of deaths?

5. 死亡人数属于哪种数据?

6. The Richter scale is used to quantify the energy produced by an earthquake. Examples of Richter scale numbers are 2.3, 4.0, 6.1, and 7.0. What kind of data are these numbers?

6. 里氏震级用于量化地震释放的能量。里氏震级数值的例子有 2.3、4.0、6.1 和 7.0。这些数字属于哪种数据?

1.4 Experimental Design and Ethics 1.4 实验设计与伦理

Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as the influence of alcohol? Questions like these are answered using randomized experiments. In this module, you will learn important aspects of experimental design. Proper study design ensures the production of reliable, accurate data.

阿司匹林能降低心脏病发作的风险吗?某品牌的肥料是否比另一种更能促进玫瑰生长?疲劳对驾驶员的危害是否与酒精的影响一样大?这类问题是通过随机实验来回答的。在本模块中,你将学习实验设计的重要方面。恰当的研究设计能够保证产生可靠、准确的数据。

The purpose of an experiment is to investigate the relationship between two variables. When one variable causes change in another, we call the first variable the explanatory variable. The affected variable is called the response variable. In a randomized experiment, the researcher manipulates values of the explanatory variable and measures the resulting changes in the response variable. The different values of the explanatory variable are called treatments. An experimental unit is a single object or individual to be measured.

实验的目的是考察两个变量之间的关系。当一个变量引起另一个变量发生改变时,我们称第一个变量为解释变量。受影响的变量称为响应变量。在随机实验中,研究者操纵解释变量的取值,并测量响应变量随之发生的变化。解释变量的不同取值称为处理。实验单元是指一个待测量的单个对象或个体。

You want to investigate the effectiveness of vitamin E in preventing disease. You recruit a group of subjects and ask them if they regularly take vitamin E. You notice that the subjects who take vitamin E exhibit better health on average than those who do not. Does this prove that vitamin E is effective in disease prevention? It does not. There are many differences between the two groups compared in addition to vitamin E consumption. People who take vitamin E regularly often take other steps to improve their health: exercise, diet, other vitamin supplements, choosing not to smoke. Any one of these factors could be influencing health. As described, this study does not prove that vitamin E is the key to disease prevention.

你想研究维生素 E 在预防疾病方面的有效性。你招募了一组受试者,询问他们是否定期服用维生素 E。你注意到,服用维生素 E 的受试者平均而言比不服用的受试者更健康。这能证明维生素 E 在疾病预防方面有效吗?并不能。除了是否服用维生素 E 之外,被比较的两个群体之间还存在许多差异。经常服用维生素 E 的人往往还会采取其他增进健康的措施:锻炼、饮食、服用其他维生素补充剂、选择不吸烟。这些因素中的任何一个都可能影响健康。正如所述,这项研究并不能证明维生素 E 是预防疾病的关键。

Additional variables that can cloud a study are called lurking variables. In order to prove that the explanatory variable is causing a change in the response variable, it is necessary to isolate the explanatory variable. The researcher must design her experiment in such a way that there is only one difference between groups being compared: the planned treatments. This is accomplished by the random assignment of experimental units to treatment groups. When subjects are assigned treatments randomly, all of the potential lurking variables are spread equally among the groups. At this point the only difference between groups is the one imposed by the researcher. Different outcomes measured in the response variable, therefore, must be a direct result of the different treatments. In this way, an experiment can prove a cause-and-effect connection between the explanatory and response variables.

可能混淆一项研究的额外变量称为潜在变量(lurking variables)。为了证明解释变量确实引起了响应变量的变化,必须隔离解释变量。研究者必须这样设计实验:在被比较的各组之间只有一种差异,即计划的处理。这是通过将实验单元随机分配到各处理组来实现的。当受试者被随机分配处理时,所有潜在的潜在变量都会在各组间均匀分散。此时,各组之间唯一的差异就是研究者所施加的那一种。因此,在响应变量上测量到的不同结果,必定是不同处理直接造成的。通过这种方式,实验可以证明解释变量与响应变量之间的因果关系。

The power of suggestion can have an important influence on the outcome of an experiment. Studies have shown that the expectation of the study participant can be as important as the actual medication. In one study of performance-enhancing drugs, researchers noted:

暗示的力量会对实验结果产生重要影响。研究表明,研究参与者的预期可能与实际药物本身同样重要。在一项关于增强表现药物的研究中,研究者指出:

*Results showed that believing one had taken the substance resulted in \[*performance*\] times almost as fast as those associated with consuming the drug itself. In contrast, taking the drug without knowledge yielded no significant performance increment.*1

*结果表明,相信自己服用了该物质所带来的\[*表现*\]时间,几乎与真正服用该药物本身所关联的时间一样快。相比之下,在不知情的情况下服用该药物并未带来显著的性能提升。*1

When participation in a study prompts a physical response from a participant, it is difficult to isolate the effects of the explanatory variable. To counter the power of suggestion, researchers set aside one treatment group as a control group. This group is given a placebo treatment–a treatment that cannot influence the response variable. The control group helps researchers balance the effects of being in an experiment with the effects of the active treatments. Of course, if you are participating in a study and you know that you are receiving a pill which contains no actual medication, then the power of suggestion is no longer a factor. Blinding in a randomized experiment preserves the power of suggestion. When a person involved in a research study is blinded, he does not know who is receiving the active treatment(s) and who is receiving the placebo treatment. A double-blind experiment is one in which both the subjects and the researchers involved with the subjects are blinded.

当参与一项研究会引发参与者的生理反应时,就很难隔离解释变量的效应。为了抵消暗示的力量,研究者将一组处理设为对照组。该组接受安慰剂处理——一种不会影响响应变量的处理。对照组帮助研究者在"参与实验本身所产生的效应"与"活性处理所产生的效应"之间取得平衡。当然,如果你参与一项研究并且知道自己服用的是不含任何实际药物的药丸,那么暗示的力量就不再是影响因素了。随机实验中的盲法保留了暗示的力量。当参与研究的人被施以盲法时,他不知道谁接受了活性处理、谁接受了安慰剂处理。双盲实验是指受试者和与受试者相关的研究者双方都被施以盲法的实验。

Problem 问题

Researchers want to investigate whether taking aspirin regularly reduces the risk of heart attack. Four hundred men between the ages of 50 and 84 are recruited as participants. The men are divided randomly into two groups: one group will take aspirin, and the other group will take a placebo. Each man takes one pill each day for three years, but he does not know whether he is taking aspirin or the placebo. At the end of the study, researchers count the number of men in each group who have had heart attacks.

研究者想要探究定期服用阿司匹林是否能降低心脏病发作的风险。他们招募了 400 名年龄在 50 至 84 岁之间的男性作为参与者。这些男性被随机分为两组:一组服用阿司匹林,另一组服用安慰剂。每名男性每天服用一粒药丸,持续三年,但他并不知道自己是服用了阿司匹林还是安慰剂。研究结束时,研究者统计每组中发生过心脏病发作的男性人数。

Identify the following values for this study: population, sample, experimental units, explanatory variable, response variable, treatments.

请为该研究确定以下各项的取值:总体、样本、实验单元、解释变量、响应变量、处理。

Problem 问题

The Smell & Taste Treatment and Research Foundation conducted a study to investigate whether smell can affect learning. Subjects completed mazes multiple times while wearing masks. They completed the pencil and paper mazes three times wearing floral-scented masks, and three times with unscented masks. Participants were assigned at random to wear the floral mask during the first three trials or during the last three trials. For each trial, researchers recorded the time it took to complete the maze and the subject’s impression of the mask’s scent: positive, negative, or neutral.

气味与味觉治疗研究基金会开展了一项研究,以探究气味是否能够影响学习。受试者戴着面罩多次完成迷宫。他们戴着花香型面罩完成铅笔纸质迷宫三次,再戴着无香型面罩完成三次。参与者被随机分配在前三次试验或后三次试验中戴花香型面罩。对于每次试验,研究者记录完成迷宫所用的时间,以及受试者对面罩气味的印象:正面、负面或中性。

1. Describe the explanatory and response variables in this study.

1. 描述本研究中的解释变量与响应变量。

2. What are the treatments?

2. 处理是什么?

3. Identify any lurking variables that could interfere with this study.

3. 找出可能干扰本研究的任何潜在变量(lurking variables)。

4. Is it possible to use blinding in this study?

4. 在本研究中是否可以使用盲法?

Problem 问题

A researcher wants to study the effects of birth order on personality. Explain why this study could not be conducted as a randomized experiment. What is the main problem in a study that cannot be designed as a randomized experiment?

一位研究者想要研究出生顺序对人格的影响。请解释为什么这项研究无法以随机实验的形式进行。在一项无法设计为随机实验的研究中,主要问题是什么?

You are concerned about the effects of texting on driving performance. Design a study to test the response time of drivers while texting and while driving only. How many seconds does it take for a driver to respond when a leading car hits the brakes?

你担心发短信对驾驶表现的影响。请设计一项研究,以测试驾驶员在发短信时与仅驾驶时的反应时间。当前车刹车时,驾驶员需要多少秒才能做出反应?

1. Describe the explanatory and response variables in the study.

1. 描述本研究中的解释变量与响应变量。

2. What are the treatments?

2. 处理是什么?

3. What should you consider when selecting participants?

3. 在选择参与者时你应该考虑什么?

4. Your research partner wants to divide participants randomly into two groups: one to drive without distraction and one to text and drive simultaneously. Is this a good idea? Why or why not?

4. 你的研究伙伴想把参与者随机分成两组:一组在无干扰的情况下驾驶,另一组同时发短信并驾驶。这是个好主意吗?为什么是或为什么不是?

5. Identify any lurking variables that could interfere with this study.

5. 找出可能干扰本研究的任何潜在变量(lurking variables)。

6. How can blinding be used in this study?

6. 在本研究中可以如何使用盲法?

Ethics 伦理

The widespread misuse and misrepresentation of statistical information often gives the field a bad name. Some say that “numbers don’t lie,” but the people who use numbers to support their claims often do.

统计信息的广泛误用与歪曲常常给这个领域带来坏名声。有人说"数字不会说谎",但那些用数字来支持自己主张的人却常常会说谎。

A recent investigation of famous social psychologist, Diederik Stapel, has led to the retraction of his articles from some of the world’s top journals including *Journal of Experimental Social Psychology, Social Psychology, Basic and Applied Social Psychology, British Journal of Social Psychology,* and the magazine *Science*. Diederik Stapel is a former professor at Tilburg University in the Netherlands. Over the past two years, an extensive investigation involving three universities where Stapel has worked concluded that the psychologist is guilty of fraud on a colossal scale. Falsified data taints over 55 papers he authored and 10 Ph.D. dissertations that he supervised.

最近对著名社会心理学家 Diederik Stapel 的调查,导致他的一些文章从世界上顶尖的期刊中被撤稿,这些期刊包括《实验社会心理学杂志》《社会心理学》《基础与应用社会心理学》《英国社会心理学杂志》以及《科学》杂志。Diederik Stapel 是荷兰蒂尔堡大学的前教授。在过去两年中,一项涉及 Stapel 曾工作过的三所大学的广泛调查得出结论:这位心理学家犯下了规模惊人的造假行为。伪造的数据玷污了他撰写的 55 篇以上论文以及他指导的 10 篇博士论文。

*Stapel did not deny that his deceit was driven by ambition. But it was more complicated than that, he told me. He insisted that he loved social psychology but had been frustrated by the messiness of experimental data, which rarely led to clear conclusions. His lifelong obsession with elegance and order, he said, led him to concoct sexy results that journals found attractive. “It was a quest for aesthetics, for beauty—instead of the truth,” he said. He described his behavior as an addiction that drove him to carry out acts of increasingly daring fraud, like a junkie seeking a bigger and better high.2*

*Stapel 并不否认他的欺骗行为是由野心驱动的。但他告诉我,事情比那更复杂。他坚称自己热爱社会心理学,却一直受困于实验数据的杂乱无章——这些数据很少能得出明确的结论。他说,他一生对优雅与秩序的痴迷,导致他编造出期刊觉得有吸引力的诱人结果。"那是对美学、对美的追求——而不是对真理的追求,"他说。他将自己的行为描述为一种成瘾,驱使着他实施越来越大胆的造假行为,就像一个瘾君子在追求更强烈、更极致的快感。2*

The committee investigating Stapel concluded that he is guilty of several practices including:

调查 Stapel 的委员会得出结论:他犯有多种行为不当,包括:

Clearly, it is never acceptable to falsify data the way this researcher did. Sometimes, however, violations of ethics are not as easy to spot.

显然,绝不可接受像这位研究者那样伪造数据。然而,有时对伦理的违反并不那么容易察觉。

Researchers have a responsibility to verify that proper methods are being followed. The report describing the investigation of Stapel’s fraud states that, “statistical flaws frequently revealed a lack of familiarity with elementary statistics.”3 Many of Stapel’s co-authors should have spotted irregularities in his data. Unfortunately, they did not know very much about statistical analysis, and they simply trusted that he was collecting and reporting data properly.

研究者有责任核实所采用的方法是否恰当。描述 Stapel 造假调查的报告指出,"统计缺陷频繁地暴露出对基础统计学的生疏。"3 Stapel 的许多合著者本应发现他数据中的异常。遗憾的是,他们对统计分析知之甚少,于是仅仅相信他在正确地收集和报告数据。

Many types of statistical fraud are difficult to spot. Some researchers simply stop collecting data once they have just enough to prove what they had hoped to prove. They don’t want to take the chance that a more extensive study would complicate their lives by producing data contradicting their hypothesis.

许多类型的统计造假都难以察觉。有些研究者一旦收集到足以证明他们原本期望证明的结论的数据,就干脆停止收集数据。他们不想冒这样的风险:一项更详尽的研究可能会产生与他们的假设相矛盾的数据,从而让他们的日子不好过。

Professional organizations, like the American Statistical Association, clearly define expectations for researchers. There are even laws in the federal code about the use of research data.

专业组织(如美国统计协会)明确规定了研究者的行为期望。联邦法典中甚至还有关于研究数据使用的法律。

When a statistical study uses human participants, as in medical studies, both ethics and the law dictate that researchers should be mindful of the safety of their research subjects. The U.S. Department of Health and Human Services oversees federal regulations of research studies with the aim of protecting participants. When a university or other research institution engages in research, it must ensure the safety of all human subjects. For this reason, research institutions establish oversight committees known as Institutional Review Boards (IRB). All planned studies must be approved in advance by the IRB. Key protections that are mandated by law include the following:

当一项统计研究使用人类参与者时(如医学研究),伦理与法律都要求研究者关注其研究对象的安全。美国卫生与公众服务部监督研究研究的联邦法规,旨在保护参与者。当大学或其他研究机构开展研究时,必须确保所有人类受试者的安全。为此,研究机构设立了被称为机构审查委员会(Institutional Review Boards,简称 IRB)的监督委员会。所有计划中的研究都必须事先获得 IRB 的批准。法律强制要求的关键保护措施包括以下各项:

These ideas may seem fundamental, but they can be very difficult to verify in practice. Is removing a participant’s name from the data record sufficient to protect privacy? Perhaps the person’s identity could be discovered from the data that remains. What happens if the study does not proceed as planned and risks arise that were not anticipated? When is informed consent really necessary? Suppose your doctor wants a blood sample to check your cholesterol level. Once the sample has been tested, you expect the lab to dispose of the remaining blood. At that point the blood becomes biological waste. Does a researcher have the right to take it for use in a study?

这些理念看似基本,但在实践中却可能很难核实。从数据记录中删除参与者的姓名是否足以保护隐私?或许仍能从剩余的数据中识别出此人的身份。如果研究未能按计划进行,并出现了未曾预料的风险,会发生什么?知情同意究竟在何时才是真正必要的?假设你的医生想要一份血样来检查你的胆固醇水平。一旦样本检测完毕,你期望实验室处置剩余的血液。此时血液就变成了生物废弃物。研究者是否有权将其取走用于某项研究?

It is important that students of statistics take time to consider the ethical questions that arise in statistical studies. How prevalent is fraud in statistical studies? You might be surprised—and disappointed. There is a website dedicated to cataloging retractions of study articles that have been proven fraudulent. A quick glance will show that the misuse of statistics is a bigger problem than most people realize.

学习统计学的学生应当花时间思考统计研究中出现的伦理问题,这一点很重要。统计研究中的造假有多普遍?你或许会感到惊讶——乃至失望。有一个网站专门收录那些已被证实造假的研究文章的撤稿信息。略览一番便会发现,统计的误用是一个比大多数人意识到的更大的问题。

Vigilance against fraud requires knowledge. Learning the basic theory of statistics will empower you to analyze statistical studies critically.

防范造假需要知识。学习统计学的基本理论将使你能够批判性地分析统计研究。

Problem 问题

Describe the unethical behavior in each example and describe how it could impact the reliability of the resulting data. Explain how the problem should be corrected.

描述每个例子中的不道德行为,并说明它可能如何影响所得数据的可靠性。解释应如何纠正该问题。

A researcher is collecting data in a community.

一位研究者正在一个社区中收集数据。

1. She selects a block where she is comfortable walking because she knows many of the people living on the street.

1. 她选择了一个自己走路感到自在的街区,因为她认识住在那条街上的许多人。

2. No one seems to be home at four houses on her route. She does not record the addresses and does not return at a later time to try to find residents at home.

2. 在她路线上的四所房子里似乎都没有人在家。她没有记录这些地址,也没有在稍后时间返回去试图找到在家的居民。

3. She skips four houses on her route because she is running late for an appointment. When she gets home, she fills in the forms by selecting random answers from other residents in the neighborhood.

3. 她跳过了路线上的四所房子,因为她赴约要迟到了。回到家后,她通过从邻里其他居民中随机选取答案来填写表格。

1.5 Data Collection Experiment 1.5 数据收集实验

Data Collection Experiment 数据收集实验

Class Time:

上课时间:

Names:

姓名:

Student Learning Outcomes

学生学习目标

Movie SurveyAsk five classmates from a different class how many movies they saw at the theater last month. Do not include rented movies.

观影调查:向另一个班级的五名同学询问他们上个月在影院观看的电影部数。不包括租借的电影。

1. Record the data.

1. 记录数据。

2. In class, randomly pick one person. On the class list, mark that person’s name. Move down four names on the class list. Mark that person’s name. Continue doing this until you have marked 12 names. You may need to go back to the start of the list. For each marked name record the five data values. You now have a total of 60 data values.

2. 在课堂上,随机选取一人。在班级名单上标记该人的姓名。沿名单向下数四个人,标记该人的姓名。继续这样做,直到标记了 12 个姓名。你可能需要回到名单开头。对每个被标记的姓名,记录五个数据值。你现在共有 60 个数据值。

3. For each name marked, record the data.

3. 对每个被标记的姓名,记录其数据。

| | | | | | | | | | | | | |-----|-----|-----|-----|-----|-----|-----|-----|-----|-----|-----|-----| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | |

| | | | | | | | | | | | | |-----|-----|-----|-----|-----|-----|-----|-----|-----|-----|-----|-----| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | |

Table 1.17

表 1.17

Order the DataComplete the two relative frequency tables below using your class data.

整理数据:使用你的班级数据完成下面的两个相对频数表。

| Number of Movies | Frequency | Relative Frequency | Cumulative Relative Frequency | |------------------|-----------|--------------------|-------------------------------| | 0 | | | | | 1 | | | | | 2 | | | | | 3 | | | | | 4 | | | | | 5 | | | | | 6 | | | | | 7+ | | | |

| 电影数量 | 频数 | 相对频数 | 累积相对频数 | |------------------|-----------|--------------------|-------------------------------| | 0 | | | | | 1 | | | | | 2 | | | | | 3 | | | | | 4 | | | | | 5 | | | | | 6 | | | | | 7+ | | | |

Table 1.18 Frequency of Number of Movies Viewed

表 1.18 观影频数

| Number of Movies | Frequency | Relative Frequency | Cumulative Relative Frequency | |------------------|-----------|--------------------|-------------------------------| | 0–1 | | | | | 2–3 | | | | | 4–5 | | | | | 6–7+ | | | |

| 电影数量 | 频数 | 相对频数 | 累积相对频数 | |------------------|-----------|--------------------|-------------------------------| | 0–1 | | | | | 2–3 | | | | | 4–5 | | | | | 6–7+ | | | |

Table 1.19 Frequency of Number of Movies Viewed

表 1.19 观影频数

1. Using the tables, find the percent of data that is at most two. Which table did you use and why?

1. 利用这些表,求至多两部电影的数据所占的百分比。你使用了哪张表,为什么?

2. Using the tables, find the percent of data that is at most three. Which table did you use and why?

2. 利用这些表,求至多三部电影的数据所占的百分比。你使用了哪张表,为什么?

3. Using the tables, find the percent of data that is more than two. Which table did you use and why?

3. 利用这些表,求多于两部电影的数据所占的百分比。你使用了哪张表,为什么?

4. Using the tables, find the percent of data that is more than three. Which table did you use and why?

4. 利用这些表,求多于三部电影的数据所占的百分比。你使用了哪张表,为什么?

Discussion Questions

讨论问题

1. Is one of the tables “more correct” than the other? Why or why not?

1. 其中一张表是否比另一张"更正确"?为什么,或为什么不?

2. In general, how could you group the data differently? Are there any advantages to either way of grouping the data?

2. 一般来说,你可以如何以不同的方式对数据进行分组?两种分组方式各自有什么优点吗?

3. Why did you switch between tables, if you did, when answering the question above?

3. 你在回答上面的问题时,如果切换了表格,为什么切换?

1.6 Sampling Experiment 1.6 抽样实验

Sampling Experiment 抽样实验

Class Time:

上课时间:

Names:

姓名:

Student Learning Outcomes

学生学习目标

In this lab, you will be asked to pick several random samples of restaurants. In each case, describe your procedure briefly, including how you might have used the random number generator, and then list the restaurants in the sample you obtained.

在本实验中,你将抽取若干餐厅的随机样本。在每种情况下,简要描述你的步骤,包括你可能如何使用随机数生成器,然后列出你所获得样本中的餐厅。

The following section contains restaurants stratified by city into columns and grouped horizontally by entree cost (clusters).

下面这一部分包含按城市分层的餐厅(列为城市),并按主菜价格(整群)横向分组。

Restaurants Stratified by City and Entree Cost

按城市与主菜价格分层的餐厅

| Entree Cost | Under \$10 | \$10 to under \$15 | \$15 to under \$20 | Over \$20 | |---------------|---------------------------------------------------------------|---------------------------------------------------------------|----------------------------------------------|-------------------------------------------| | San Jose | El Abuelo Taq, Pasta Mia, Emma’s Express, Bamboo Hut | Emperor’s Guard, Creekside Inn | Agenda, Gervais, Miro’s | Blake’s, Eulipia, Hayes Mansion, Germania | | Palo Alto | Senor Taco, Olive Garden, Taxi’s | Ming’s, P.A. Joe’s, Stickney’s | Scott’s Seafood, Poolside Grill, Fish Market | Sundance Mine, Maddalena’s, Spago’s | | Los Gatos | Mary’s Patio, Mount Everest, Sweet Pea’s, Andele Taqueria | Lindsey’s, Willow Street | Toll House | Charter House, La Maison Du Cafe | | Mountain View | Maharaja, New Ma’s, Thai-Rific, Garden Fresh | Amber Indian, La Fiesta, Fiesta del Mar, Dawit | Austin’s, Shiva’s, Mazeh | Le Petit Bistro | | Cupertino | Hobees, Hung Fu, Samrat, Panda Express | Santa Barb. Grill, Mand. Gourmet, Bombay Oven, Kathmandu West | Fontana’s, Blue Pheasant | Hamasushi, Helios | | Sunnyvale | Chekijababi, Taj India, Full Throttle, Tia Juana, Lemon Grass | Pacific Fresh, Charley Brown’s, Cafe Cameroon, Faz, Aruba’s | Lion & Compass, The Palace, Beau Sejour | | | Santa Clara | Rangoli, Armadillo Willy’s, Thai Pepper, Pasand | Arthur’s, Katie’s Cafe, Pedro’s, La Galleria | Birk’s, Truya Sushi, Valley Plaza | Lakeside, Mariani’s |

| 主菜价格 | 低于\$10 | \$10至低于\$15 | \$15至低于\$20 | 高于\$20 | |---------------|---------------------------------------------------------------|---------------------------------------------------------------|----------------------------------------------|-------------------------------------------| | San Jose | El Abuelo Taq, Pasta Mia, Emma’s Express, Bamboo Hut | Emperor’s Guard, Creekside Inn | Agenda, Gervais, Miro’s | Blake’s, Eulipia, Hayes Mansion, Germania | | Palo Alto | Senor Taco, Olive Garden, Taxi’s | Ming’s, P.A. Joe’s, Stickney’s | Scott’s Seafood, Poolside Grill, Fish Market | Sundance Mine, Maddalena’s, Spago’s | | Los Gatos | Mary’s Patio, Mount Everest, Sweet Pea’s, Andele Taqueria | Lindsey’s, Willow Street | Toll House | Charter House, La Maison Du Cafe | | Mountain View | Maharaja, New Ma’s, Thai-Rific, Garden Fresh | Amber Indian, La Fiesta, Fiesta del Mar, Dawit | Austin’s, Shiva’s, Mazeh | Le Petit Bistro | | Cupertino | Hobees, Hung Fu, Samrat, Panda Express | Santa Barb. Grill, Mand. Gourmet, Bombay Oven, Kathmandu West | Fontana’s, Blue Pheasant | Hamasushi, Helios | | Sunnyvale | Chekijababi, Taj India, Full Throttle, Tia Juana, Lemon Grass | Pacific Fresh, Charley Brown’s, Cafe Cameroon, Faz, Aruba’s | Lion & Compass, The Palace, Beau Sejour | | | Santa Clara | Rangoli, Armadillo Willy’s, Thai Pepper, Pasand | Arthur’s, Katie’s Cafe, Pedro’s, La Galleria | Birk’s, Truya Sushi, Valley Plaza | Lakeside, Mariani’s |

Table 1.20 Restaurants Used In Sample

表 1.20 抽样所用餐厅

A Simple Random SamplePick a simple random sample of 15 restaurants.

简单随机样本:抽取 15 家餐厅的简单随机样本

1. Describe your procedure.

1. 描述你的步骤。

2. Complete the table with your sample.

2. 用你的样本填写表格。

| | | | |--------------------------|---------------------------|---------------------------| | 1\. \_\_\_\_\_\_\_\_\_\_ | 6\. \_\_\_\_\_\_\_\_\_\_ | 11\. \_\_\_\_\_\_\_\_\_\_ | | 2\. \_\_\_\_\_\_\_\_\_\_ | 7\. \_\_\_\_\_\_\_\_\_\_ | 12\. \_\_\_\_\_\_\_\_\_\_ | | 3\. \_\_\_\_\_\_\_\_\_\_ | 8\. \_\_\_\_\_\_\_\_\_\_ | 13\. \_\_\_\_\_\_\_\_\_\_ | | 4\. \_\_\_\_\_\_\_\_\_\_ | 9\. \_\_\_\_\_\_\_\_\_\_ | 14\. \_\_\_\_\_\_\_\_\_\_ | | 5\. \_\_\_\_\_\_\_\_\_\_ | 10\. \_\_\_\_\_\_\_\_\_\_ | 15\. \_\_\_\_\_\_\_\_\_\_ |

| | | | |--------------------------|---------------------------|---------------------------| | 1\. \_\_\_\_\_\_\_\_\_\_ | 6\. \_\_\_\_\_\_\_\_\_\_ | 11\. \_\_\_\_\_\_\_\_\_\_ | | 2\. \_\_\_\_\_\_\_\_\_\_ | 7\. \_\_\_\_\_\_\_\_\_\_ | 12\. \_\_\_\_\_\_\_\_\_\_ | | 3\. \_\_\_\_\_\_\_\_\_\_ | 8\. \_\_\_\_\_\_\_\_\_\_ | 13\. \_\_\_\_\_\_\_\_\_\_ | | 4\. \_\_\_\_\_\_\_\_\_\_ | 9\. \_\_\_\_\_\_\_\_\_\_ | 14\. \_\_\_\_\_\_\_\_\_\_ | | 5\. \_\_\_\_\_\_\_\_\_\_ | 10\. \_\_\_\_\_\_\_\_\_\_ | 15\. \_\_\_\_\_\_\_\_\_\_ |

Table 1.21

表 1.21

A Systematic SamplePick a systematic sample of 15 restaurants.

系统抽样:抽取 15 家餐厅的系统抽样样本。

1. Describe your procedure.

1. 描述你的步骤。

2. Complete the table with your sample.

2. 用你的样本填写表格。

| | | | |--------------------------|---------------------------|---------------------------| | 1\. \_\_\_\_\_\_\_\_\_\_ | 6\. \_\_\_\_\_\_\_\_\_\_ | 11\. \_\_\_\_\_\_\_\_\_\_ | | 2\. \_\_\_\_\_\_\_\_\_\_ | 7\. \_\_\_\_\_\_\_\_\_\_ | 12\. \_\_\_\_\_\_\_\_\_\_ | | 3\. \_\_\_\_\_\_\_\_\_\_ | 8\. \_\_\_\_\_\_\_\_\_\_ | 13\. \_\_\_\_\_\_\_\_\_\_ | | 4\. \_\_\_\_\_\_\_\_\_\_ | 9\. \_\_\_\_\_\_\_\_\_\_ | 14\. \_\_\_\_\_\_\_\_\_\_ | | 5\. \_\_\_\_\_\_\_\_\_\_ | 10\. \_\_\_\_\_\_\_\_\_\_ | 15\. \_\_\_\_\_\_\_\_\_\_ |

| | | | |--------------------------|---------------------------|---------------------------| | 1\. \_\_\_\_\_\_\_\_\_\_ | 6\. \_\_\_\_\_\_\_\_\_\_ | 11\. \_\_\_\_\_\_\_\_\_\_ | | 2\. \_\_\_\_\_\_\_\_\_\_ | 7\. \_\_\_\_\_\_\_\_\_\_ | 12\. \_\_\_\_\_\_\_\_\_\_ | | 3\. \_\_\_\_\_\_\_\_\_\_ | 8\. \_\_\_\_\_\_\_\_\_\_ | 13\. \_\_\_\_\_\_\_\_\_\_ | | 4\. \_\_\_\_\_\_\_\_\_\_ | 9\. \_\_\_\_\_\_\_\_\_\_ | 14\. \_\_\_\_\_\_\_\_\_\_ | | 5\. \_\_\_\_\_\_\_\_\_\_ | 10\. \_\_\_\_\_\_\_\_\_\_ | 15\. \_\_\_\_\_\_\_\_\_\_ |

Table 1.22

表 1.22

A Stratified SamplePick a stratified sample, by city, of 20 restaurants. Use 25% of the restaurants from each stratum. Round to the nearest whole number.

分层样本:按城市抽取 20 家餐厅的分层样本。每个层使用 25% 的餐厅。四舍五入到最接近的整数。

1. Describe your procedure.

1. 描述你的步骤。

2. Complete the table with your sample.

2. 用你的样本填写表格。

| | | | | |--------------------------|---------------------------|---------------------------|---------------------------| | 1\. \_\_\_\_\_\_\_\_\_\_ | 6\. \_\_\_\_\_\_\_\_\_\_ | 11\. \_\_\_\_\_\_\_\_\_\_ | 16\. \_\_\_\_\_\_\_\_\_\_ | | 2\. \_\_\_\_\_\_\_\_\_\_ | 7\. \_\_\_\_\_\_\_\_\_\_ | 12\. \_\_\_\_\_\_\_\_\_\_ | 17\. \_\_\_\_\_\_\_\_\_\_ | | 3\. \_\_\_\_\_\_\_\_\_\_ | 8\. \_\_\_\_\_\_\_\_\_\_ | 13\. \_\_\_\_\_\_\_\_\_\_ | 18\. \_\_\_\_\_\_\_\_\_\_ | | 4\. \_\_\_\_\_\_\_\_\_\_ | 9\. \_\_\_\_\_\_\_\_\_\_ | 14\. \_\_\_\_\_\_\_\_\_\_ | 19\. \_\_\_\_\_\_\_\_\_\_ | | 5\. \_\_\_\_\_\_\_\_\_\_ | 10\. \_\_\_\_\_\_\_\_\_\_ | 15\. \_\_\_\_\_\_\_\_\_\_ | 20\. \_\_\_\_\_\_\_\_\_\_ |

| | | | | |--------------------------|---------------------------|---------------------------|---------------------------| | 1\. \_\_\_\_\_\_\_\_\_\_ | 6\. \_\_\_\_\_\_\_\_\_\_ | 11\. \_\_\_\_\_\_\_\_\_\_ | 16\. \_\_\_\_\_\_\_\_\_\_ | | 2\. \_\_\_\_\_\_\_\_\_\_ | 7\. \_\_\_\_\_\_\_\_\_\_ | 12\. \_\_\_\_\_\_\_\_\_\_ | 17\. \_\_\_\_\_\_\_\_\_\_ | | 3\. \_\_\_\_\_\_\_\_\_\_ | 8\. \_\_\_\_\_\_\_\_\_\_ | 13\. \_\_\_\_\_\_\_\_\_\_ | 18\. \_\_\_\_\_\_\_\_\_\_ | | 4\. \_\_\_\_\_\_\_\_\_\_ | 9\. \_\_\_\_\_\_\_\_\_\_ | 14\. \_\_\_\_\_\_\_\_\_\_ | 19\. \_\_\_\_\_\_\_\_\_\_ | | 5\. \_\_\_\_\_\_\_\_\_\_ | 10\. \_\_\_\_\_\_\_\_\_\_ | 15\. \_\_\_\_\_\_\_\_\_\_ | 20\. \_\_\_\_\_\_\_\_\_\_ |

Table 1.23

表 1.23

A Stratified SamplePick a stratified sample, by entree cost, of 21 restaurants. Use 25% of the restaurants from each stratum. Round to the nearest whole number.

分层样本:按主菜价格抽取 21 家餐厅的分层样本。每个层使用 25% 的餐厅。四舍五入到最接近的整数。

1. Describe your procedure.

1. 描述你的步骤。

2. Complete the table with your sample.

2. 用你的样本填写表格。

| | | | | |--------------------------|---------------------------|---------------------------|---------------------------| | 1\. \_\_\_\_\_\_\_\_\_\_ | 6\. \_\_\_\_\_\_\_\_\_\_ | 11\. \_\_\_\_\_\_\_\_\_\_ | 16\. \_\_\_\_\_\_\_\_\_\_ | | 2\. \_\_\_\_\_\_\_\_\_\_ | 7\. \_\_\_\_\_\_\_\_\_\_ | 12\. \_\_\_\_\_\_\_\_\_\_ | 17\. \_\_\_\_\_\_\_\_\_\_ | | 3\. \_\_\_\_\_\_\_\_\_\_ | 8\. \_\_\_\_\_\_\_\_\_\_ | 13\. \_\_\_\_\_\_\_\_\_\_ | 18\. \_\_\_\_\_\_\_\_\_\_ | | 4\. \_\_\_\_\_\_\_\_\_\_ | 9\. \_\_\_\_\_\_\_\_\_\_ | 14\. \_\_\_\_\_\_\_\_\_\_ | 19\. \_\_\_\_\_\_\_\_\_\_ | | 5\. \_\_\_\_\_\_\_\_\_\_ | 10\. \_\_\_\_\_\_\_\_\_\_ | 15\. \_\_\_\_\_\_\_\_\_\_ | 20\. \_\_\_\_\_\_\_\_\_\_ | | | | | 21\. \_\_\_\_\_\_\_\_\_\_ |

| | | | | |--------------------------|---------------------------|---------------------------|---------------------------| | 1\. \_\_\_\_\_\_\_\_\_\_ | 6\. \_\_\_\_\_\_\_\_\_\_ | 11\. \_\_\_\_\_\_\_\_\_\_ | 16\. \_\_\_\_\_\_\_\_\_\_ | | 2\. \_\_\_\_\_\_\_\_\_\_ | 7\. \_\_\_\_\_\_\_\_\_\_ | 12\. \_\_\_\_\_\_\_\_\_\_ | 17\. \_\_\_\_\_\_\_\_\_\_ | | 3\. \_\_\_\_\_\_\_\_\_\_ | 8\. \_\_\_\_\_\_\_\_\_\_ | 13\. \_\_\_\_\_\_\_\_\_\_ | 18\. \_\_\_\_\_\_\_\_\_\_ | | 4\. \_\_\_\_\_\_\_\_\_\_ | 9\. \_\_\_\_\_\_\_\_\_\_ | 14\. \_\_\_\_\_\_\_\_\_\_ | 19\. \_\_\_\_\_\_\_\_\_\_ | | 5\. \_\_\_\_\_\_\_\_\_\_ | 10\. \_\_\_\_\_\_\_\_\_\_ | 15\. \_\_\_\_\_\_\_\_\_\_ | 20\. \_\_\_\_\_\_\_\_\_\_ | | | | | 21\. \_\_\_\_\_\_\_\_\_\_ |

Table 1.24

表 1.24

A Cluster SamplePick a cluster sample of restaurants from two cities. The number of restaurants will vary.

整群样本:从两座城市抽取餐厅的整群样本。餐厅数量会有所不同。

1. Describe your procedure.

1. 描述你的步骤。

2. Complete the table with your sample.

2. 用你的样本填写表格。

| | | | | | |----------------------|-----------------------|-----------------------|-----------------------|-----------------------| | 1\. \_\_\_\_\_\_\_\_ | 6\. \_\_\_\_\_\_\_\_ | 11\. \_\_\_\_\_\_\_\_ | 16\. \_\_\_\_\_\_\_\_ | 21\. \_\_\_\_\_\_\_\_ | | 2\. \_\_\_\_\_\_\_\_ | 7\. \_\_\_\_\_\_\_\_ | 12\. \_\_\_\_\_\_\_\_ | 17\. \_\_\_\_\_\_\_\_ | 22\. \_\_\_\_\_\_\_\_ | | 3\. \_\_\_\_\_\_\_\_ | 8\. \_\_\_\_\_\_\_\_ | 13\. \_\_\_\_\_\_\_\_ | 18\. \_\_\_\_\_\_\_\_ | 23\. \_\_\_\_\_\_\_\_ | | 4\. \_\_\_\_\_\_\_\_ | 9\. \_\_\_\_\_\_\_\_ | 14\. \_\_\_\_\_\_\_\_ | 19\. \_\_\_\_\_\_\_\_ | 24\. \_\_\_\_\_\_\_\_ | | 5\. \_\_\_\_\_\_\_\_ | 10\. \_\_\_\_\_\_\_\_ | 15\. \_\_\_\_\_\_\_\_ | 20\. \_\_\_\_\_\_\_\_ | 25\. \_\_\_\_\_\_\_\_ |

| | | | | | |----------------------|-----------------------|-----------------------|-----------------------|-----------------------| | 1\. \_\_\_\_\_\_\_\_ | 6\. \_\_\_\_\_\_\_\_ | 11\. \_\_\_\_\_\_\_\_ | 16\. \_\_\_\_\_\_\_\_ | 21\. \_\_\_\_\_\_\_\_ | | 2\. \_\_\_\_\_\_\_\_ | 7\. \_\_\_\_\_\_\_\_ | 12\. \_\_\_\_\_\_\_\_ | 17\. \_\_\_\_\_\_\_\_ | 22\. \_\_\_\_\_\_\_\_ | | 3\. \_\_\_\_\_\_\_\_ | 8\. \_\_\_\_\_\_\_\_ | 13\. \_\_\_\_\_\_\_\_ | 18\. \_\_\_\_\_\_\_\_ | 23\. \_\_\_\_\_\_\_\_ | | 4\. \_\_\_\_\_\_\_\_ | 9\. \_\_\_\_\_\_\_\_ | 14\. \_\_\_\_\_\_\_\_ | 19\. \_\_\_\_\_\_\_\_ | 24\. \_\_\_\_\_\_\_\_ | | 5\. \_\_\_\_\_\_\_\_ | 10\. \_\_\_\_\_\_\_\_ | 15\. \_\_\_\_\_\_\_\_ | 20\. \_\_\_\_\_\_\_\_ | 25\. \_\_\_\_\_\_\_\_ |

Table 1.25

表 1.25

Key Terms 关键术语

Average

平均数

also called mean; a number that describes the central tendency of the data

又称均值;描述数据集中趋势的一个数值

Blinding

盲法

not telling participants which treatment a subject is receiving

不告知参与者某受试者所接受的处理

Categorical Variable

分类变量

variables that take on values that are names or labels

取值为名称或标签的变量

Cluster Sampling

整群抽样

a method for selecting a random sample and dividing the population into groups (clusters); use simple random sampling to select a set of clusters. Every individual in the chosen clusters is included in the sample.

一种选取随机样本的方法,将总体划分为若干组(群);使用简单随机抽样选取若干群。所选群中的每一个个体都纳入样本。

Continuous Random Variable

连续随机变量

a random variable (RV) whose outcomes are measured; the height of trees in the forest is a continuous RV.

结果通过测量得到的随机变量(RV);森林中树木的高度就是一个连续随机变量。

Control Group

对照组

a group in a randomized experiment that receives an inactive treatment but is otherwise managed exactly as the other groups

在随机化实验中接受无效处理,但在其他方面与其他组受到完全相同管理的组

Convenience Sampling

方便抽样

a nonrandom method of selecting a sample; this method selects individuals that are easily accessible and may result in biased data.

一种非随机的样本选取方法;该方法选取容易获得的个体,可能产生有偏的数据。

Cumulative Relative Frequency

累积相对频数

The term applies to an ordered set of observations from smallest to largest. The cumulative relative frequency is the sum of the relative frequencies for all values that are less than or equal to the given value.

该术语适用于从小到大排序的一组观测值。累积相对频数是所有小于或等于给定值的数据的相对频数之和。

Data

数据

a set of observations (a set of possible outcomes); most data can be put into two groups: qualitative (an attribute whose value is indicated by a label) or quantitative (an attribute whose value is indicated by a number). Quantitative data can be separated into two subgroups: discrete and continuous. Data is discrete if it is the result of counting (such as the number of students of a given ethnic group in a class or the number of books on a shelf). Data is continuous if it is the result of measuring (such as distance traveled or weight of luggage)

一组观测值(一组可能的结果);多数数据可分为两类:定性数据(其值由标签表示的属性)或定量数据(其值由数字表示的属性)。定量数据可进一步分为两个子类:离散数据连续数据。若数据是计数结果(例如一个班级中某一种族的学生人数,或书架上书的数量),则为离散数据。若数据是测量所得(例如行驶距离或行李重量),则为连续数据

Discrete Random Variable

离散随机变量

a random variable (RV) whose outcomes are counted

结果通过计数得到的随机变量(RV)

Double-blinding

双盲

the act of blinding both the subjects of an experiment and the researchers who work with the subjects

对实验受试者和与受试者合作的研究者双方均实施盲法的做法

Experimental Unit

实验单元

any individual or object to be measured

任何待测量的个体或对象

Explanatory Variable

解释变量

the independent variable in an experiment; the value controlled by researchers

实验中的自变量;由研究者控制其取值的变量

Frequency

频数

the number of times a value of the data occurs

某一数据值出现的次数

Informed Consent

知情同意

Any human subject in a research study must be cognizant of any risks or costs associated with the study. The subject has the right to know the nature of the treatments included in the study, their potential risks, and their potential benefits. Consent must be given freely by an informed, fit participant.

研究中的任何人类受试者都必须了解与本研究相关的任何风险或代价。受试者有权知道研究所包含的处理的性质、其潜在风险与潜在收益。同意必须由知情且合格的受试者自由地给出。

Institutional Review Board

机构审查委员会

a committee tasked with oversight of research programs that involve human subjects

负责监督涉及人类受试者的研究项目的委员会

Lurking Variable

潜在变量

a variable that has an effect on a study even though it is neither an explanatory variable nor a response variable

对研究有影响,但既不是解释变量也不是响应变量的变量

Nonsampling Error

非抽样误差

an issue that affects the reliability of sampling data other than natural variation; it includes a variety of human errors including poor study design, biased sampling methods, inaccurate information provided by study participants, data entry errors, and poor analysis.

除自然变异外,影响抽样数据可靠性的问题;它包括多种人为错误,如研究设计不佳、有偏的抽样方法、受试者提供的信息不准确、数据录入错误以及分析不当。

Numerical Variable

数值变量

variables that take on values that are indicated by numbers

取值由数字表示的变量

Parameter

参数

a number that is used to represent a population characteristic and that generally cannot be determined easily

用于表示总体特征、通常不易确定的数值

Placebo

安慰剂

an inactive treatment that has no real effect on the explanatory variable

对解释变量无真实作用的无效处理

Population

总体

all individuals, objects, or measurements whose properties are being studied

其属性正被研究的所有个体、对象或测量值

Probability

概率

a number between zero and one, inclusive, that gives the likelihood that a specific event will occur

介于零与一(含端点)之间的数,表示某一特定事件发生的可能性

Proportion

比例

the number of successes divided by the total number in the sample

成功次数除以样本总数所得之值

Qualitative Data

定性数据

See Data.

见"数据"。

Quantitative Data

定量数据

See Data.

见"数据"。

Random Assignment

随机分配

the act of organizing experimental units into treatment groups using random methods

使用随机方法将实验单元分配到各处理组的行为

Random Sampling

随机抽样

a method of selecting a sample that gives every member of the population an equal chance of being selected.

一种选取样本的方法,使总体中每个成员都有同等机会被选中。

Relative Frequency

相对频数

the ratio of the number of times a value of the data occurs in the set of all outcomes to the number of all outcomes to the total number of outcomes

数据某取值在全部结果中出现的次数,与全部结果总数之比

Representative Sample

代表性样本

a subset of the population that has the same characteristics as the population

具有与总体相同特征的总体子集

Response Variable

响应变量

the dependent variable in an experiment; the value that is measured for change at the end of an experiment

实验中的因变量;在实验结束时测量其变化的取值

Sample

样本

a subset of the population studied

所研究总体的一个子集

Sampling Bias

抽样偏倚

not all members of the population are equally likely to be selected

并非总体中所有成员都有同等机会被选中

Sampling Error

抽样误差

the natural variation that results from selecting a sample to represent a larger population; this variation decreases as the sample size increases, so selecting larger samples reduces sampling error.

因选取样本代表更大总体而产生的自然变异;这种变异随样本量增大而减小,因此选取更大的样本可降低抽样误差。

Sampling with Replacement

有放回抽样

Once a member of the population is selected for inclusion in a sample, that member is returned to the population for the selection of the next individual.

一旦总体中某个成员被选中纳入样本,该成员便被放回总体,以供选取下一个个体。

Sampling without Replacement

无放回抽样

A member of the population may be chosen for inclusion in a sample only once. If chosen, the member is not returned to the population before the next selection.

总体中的成员至多只能被选中纳入样本一次。一旦被选中,在下一次选取之前该成员不会放回总体。

Simple Random Sampling

简单随机抽样

a straightforward method for selecting a random sample; give each member of the population a number. Use a random number generator to select a set of labels. These randomly selected labels identify the members of your sample.

一种直接选取随机样本的方法:给总体中每个成员编一个号,用随机数生成器选取一组编号,这些随机选中的编号即标识了样本的成员。

Statistic

统计量

a numerical characteristic of the sample; a statistic estimates the corresponding population parameter.

样本的数值特征;统计量估计相应的总体参数。

Stratified Sampling

分层抽样

a method for selecting a random sample used to ensure that subgroups of the population are represented adequately; divide the population into groups (strata). Use simple random sampling to identify a proportionate number of individuals from each stratum.

一种用于确保总体各子组得到充分代表的随机样本选取方法:将总体划分为若干组(层),对每一层使用简单随机抽样抽取与之成比例数量的个体。

Systematic Sampling

系统抽样

a method for selecting a random sample; list the members of the population. Use simple random sampling to select a starting point in the population. Let k = (number of individuals in the population)/(number of individuals needed in the sample). Choose every kth individual in the list starting with the one that was randomly selected. If necessary, return to the beginning of the population list to complete your sample.

一种选取随机样本的方法:列出总体成员,用简单随机抽样在总体中选出一个起点。令 k =(总体中个体数)/(样本所需个体数)。从随机选中的个体开始,在列表中每隔第 k 个个体选取一个。如有必要,可回到总体列表的开头以完成抽样。

Treatments

处理

different values or components of the explanatory variable applied in an experiment

实验中施加的解释变量的不同取值或组成部分

Variable

变量

a characteristic of interest for each person or object in a population

总体中每个个体或对象所关注的属性

Chapter Review 本章回顾

1.1 Definitions of Statistics, Probability, and Key Terms 1.1 统计学、概率与关键术语的定义

The mathematical theory of statistics is easier to learn when you know the language. This module presents important terms that will be used throughout the text.

当你掌握了相关语言后,统计学的数学理论更容易学习。本模块介绍全书中将会用到的重要术语。

1.2 Data, Sampling, and Variation in Data and Sampling 1.2 数据、抽样与数据及抽样中的变异

Data are individual items of information that come from a population or sample. Data may be classified as qualitative (categorical), quantitative continuous, or quantitative discrete.

数据来自总体或样本的单项信息。数据可分为定性(分类)数据、定量连续数据或定量离散数据。

Because it is not practical to measure the entire population in a study, researchers use samples to represent the population. A random sample is a representative group from the population chosen by using a method that gives each individual in the population an equal chance of being included in the sample. Random sampling methods include simple random sampling, stratified sampling, cluster sampling, and systematic sampling. Convenience sampling is a nonrandom method of choosing a sample that often produces biased data.

由于在实际研究中测量整个总体并不现实,研究者使用样本来代表总体。随机样本是总体中具有代表性的一个组,其选取方法使总体中每个个体都有同等机会被纳入样本。随机抽样方法包括简单随机抽样、分层抽样、整群抽样与系统抽样。方便抽样是一种非随机的选样方法,常常产生有偏的数据。

Samples that contain different individuals result in different data. This is true even when the samples are well-chosen and representative of the population. When properly selected, larger samples model the population more closely than smaller samples. There are many different potential problems that can affect the reliability of a sample. Statistical data needs to be critically analyzed, not simply accepted.

包含不同个体的样本会得到不同的数据。即使样本选取得当并能代表总体,这一点也仍然成立。当选取恰当时,较大的样本比较小样本更能贴近地刻画总体。存在许多可能影响样本可靠性的潜在问题。统计资料需要批判性地分析,而不能简单地照单全收。

1.3 Frequency, Frequency Tables, and Levels of Measurement 1.3 频数、频数表与测量尺度

Some calculations generate numbers that are artificially precise. It is not necessary to report a value to eight decimal places when the measures that generated that value were only accurate to the nearest tenth. Round off your final answer to one more decimal place than was present in the original data. This means that if you have data measured to the nearest tenth of a unit, report the final statistic to the nearest hundredth.

某些计算会产生人为精确的数字。当生成该值的测量本身只精确到十分位时,没有必要把结果报告到小数点后八位。将你的最终答案舍入到比原始数据多一位小数。也就是说,如果你有精确到十分位的数据,则把最终统计量报告到百分位。

In addition to rounding your answers, you can measure your data using the following four levels of measurement.

除了对答案进行舍入之外,你还可以用以下四种测量尺度来度量数据。

When organizing data, it is important to know how many times a value appears. How many statistics students study five hours or more for an exam? What percent of families on our block own two pets? Frequency, relative frequency, and cumulative relative frequency are measures that answer questions like these.

在整理数据时,了解某个数值出现多少次很重要。有多少名统计学专业的学生为考试学习五小时或更久?我们街区上有百分之多少的家庭养了两只宠物?频数、相对频数与累积相对频数正是用来回答此类问题的度量。

1.4 Experimental Design and Ethics 1.4 实验设计与伦理

A poorly designed study will not produce reliable data. There are certain key components that must be included in every experiment. To eliminate lurking variables, subjects must be assigned randomly to different treatment groups. One of the groups must act as a control group, demonstrating what happens when the active treatment is not applied. Participants in the control group receive a placebo treatment that looks exactly like the active treatments but cannot influence the response variable. To preserve the integrity of the placebo, both researchers and subjects may be blinded. When a study is designed properly, the only difference between treatment groups is the one imposed by the researcher. Therefore, when groups respond differently to different treatments, the difference must be due to the influence of the explanatory variable.

设计拙劣的研究不会产生可靠的数据。每个实验都必须包含若干关键要素。为消除潜在变量,受试对象必须被随机分配到不同的处理组。其中一组必须作为对照组,以显示未施加主动处理时会发生什么。对照组的参与者接受一种安慰剂处理,其在外观上与主动处理完全相同,但不会影响因变量。为维护安慰剂的完整性,研究人员与受试对象都可能需要被施盲。当研究设计恰当时,各处理组之间唯一的差异就是研究者所施加的那一个。因此,当各组对不同处理作出不同反应时,该差异必定是由于解释变量的影响所致。

“An ethics problem arises when you are considering an action that benefits you or some cause you support, hurts or reduces benefits to others, and violates some rule.” (Andrew Gelman, “Open Data and Open Methods,” Ethics and Statistics, http://www.stat.columbia.edu/~gelman/research/published/ChanceEthics1.pdf (accessed May 1, 2013).) Ethical violations in statistics are not always easy to spot. Professional associations and federal agencies post guidelines for proper conduct. It is important that you learn basic statistical procedures so that you can recognize proper data analysis.

“当你考虑采取一项使你本人或你所支持的某项事业获益、却损害或减少他人利益、并且违反某种规则的行动时,就会出现伦理问题。”(Andrew Gelman,《开放数据与开放方法》,《伦理与统计》,http://www.stat.columbia.edu/~gelman/research/published/ChanceEthics1.pdf(2013 年 5 月 1 日访问)。)统计学中的伦理违规并非总是一眼可辨。专业协会与联邦机构会发布规范行为的准则。学习基本的统计方法十分重要,这样你才能识别恰当的数据分析。

Practice 练习

1.1 Definitions of Statistics, Probability, and Key Terms 1.1 统计学、概率与关键术语的定义

*Use the following information to answer the next five exercises.* Studies are often done by pharmaceutical companies to determine the effectiveness of a treatment program. Suppose that a new AIDS antibody drug is currently under study. It is given to patients once the AIDS symptoms have revealed themselves. Of interest is the average (mean) length of time in months patients live once they start the treatment. Two researchers each follow a different set of 40 patients with AIDS from the start of treatment until their deaths. The following data (in months) are collected.

(用以下信息回答接下来的五道练习。)制药公司常开展研究以确定某个治疗方案的疗效。假设目前一种新型艾滋病抗体药物正在研究中。患者在艾滋病症状显现后开始用药。研究者关注的是患者开始治疗后存活时间(以月计)的平均数。两位研究者各自追踪一组不同的 40 名艾滋病患者,从治疗开始直至其死亡。收集到如下数据(单位:月)。

Researcher A:3; 4; 11; 15; 16; 17; 22; 44; 37; 16; 14; 24; 25; 15; 26; 27; 33; 29; 35; 44; 13; 21; 22; 10; 12; 8; 40; 32; 26; 27; 31; 34; 29; 17; 8; 24; 18; 47; 33; 34

研究者 A:3; 4; 11; 15; 16; 17; 22; 44; 37; 16; 14; 24; 25; 15; 26; 27; 33; 29; 35; 44; 13; 21; 22; 10; 12; 8; 40; 32; 26; 27; 31; 34; 29; 17; 8; 24; 18; 47; 33; 34

Researcher B:3; 14; 11; 5; 16; 17; 28; 41; 31; 18; 14; 14; 26; 25; 21; 22; 31; 2; 35; 44; 23; 21; 21; 16; 12; 18; 41; 22; 16; 25; 33; 34; 29; 13; 18; 24; 23; 42; 33; 29

研究者 B:3; 14; 11; 5; 16; 17; 28; 41; 31; 18; 14; 14; 26; 25; 21; 22; 31; 2; 35; 44; 23; 21; 21; 16; 12; 18; 41; 22; 16; 25; 33; 34; 29; 13; 18; 24; 23; 42; 33; 29

Determine what the key terms refer to in the example for Researcher A.

确定在上述研究者 A 的例子中,各关键术语所指的对象。

1.

1.

population

总体

2\.

2.

sample

样本

3.

3.

parameter

参数

4\.

4.

statistic

统计量

5.

5.

variable

变量

1.2 Data, Sampling, and Variation in Data and Sampling 1.2 数据、抽样与数据及抽样中的变异

6\.

6.

“Number of times per week” is what type of data?

“每周次数”是何种类型的数据?

a\. qualitative (categorical); b. quantitative discrete; c. quantitative continuous

a. 定性(分类)数据;b. 定量离散数据;c. 定量连续数据

*Use the following information to answer the next four exercises:* A study was done to determine the age, number of times per week, and the duration (amount of time) of residents using a local park in San Antonio, Texas. The first house in the neighborhood around the park was selected randomly, and then the resident of every eighth house in the neighborhood around the park was interviewed.

(用以下信息回答接下来的四道练习。)一项研究旨在了解得克萨斯州圣安东尼奥市当地一个公园使用者的情况,包括其年龄、每周使用次数以及每次使用时长(时间量)。先随机选取公园周边社区的第一户住宅,然后访问该社区中每隔八户的住户。

7.

7.

The sampling method was

抽样方法是

a\. simple random; b. systematic; c. stratified; d. cluster

a. 简单随机;b. 系统抽样;c. 分层;d. 整群

8\.

8.

“Duration (amount of time)” is what type of data?

“时长(时间量)”是何种类型的数据?

a\. qualitative (categorical); b. quantitative discrete; c. quantitative continuous

a. 定性(分类)数据;b. 定量离散数据;c. 定量连续数据

9.

9.

The colors of the houses around the park are what kind of data?

公园周边房屋的颜色是何种数据?

a\. qualitative (categorical); b. quantitative discrete; c. quantitative continuous

a. 定性(分类)数据;b. 定量离散数据;c. 定量连续数据

10\.

10.

The population is \_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_

总体是_____________

11.

11.

Table 1.26 contains the total number of deaths worldwide as a result of earthquakes from 2000 to 2012.

表 1.26 列出了 2000 至 2012 年间全球因地震死亡的总人数。
Table 1.26
YearTotal Number of Deaths
2000231
200121,357
200211,685
200333,819
2004228,802
200588,003
20066,605
2007712
200888,011
20091,790
2010320,120
201121,953
2012768
Total823,856
表 1.26
年份死亡总人数
2000231
200121,357
200211,685
200333,819
2004228,802
200588,003
20066,605
2007712
200888,011
20091,790
2010320,120
201121,953
2012768
合计823,856

Use Table 1.26 to answer the following questions.

用表 1.26 回答以下问题。

1. What is the proportion of deaths between 2007 and 2012?

1. 2007 至 2012 年间的死亡人数所占比例是多少?

2. What percent of deaths occurred before 2001?

2. 2001 年之前的死亡人数占百分之多少?

3. What is the percent of deaths that occurred in 2003 or after 2010?

3. 发生在 2003 年或 2010 年及以后的死亡人数占百分之多少?

4. What is the fraction of deaths that happened before 2012?

4. 2012 年之前发生的死亡人数占几分之几?

5. What kind of data is the number of deaths?

5. 死亡人数是何种类型的数据?

6. Earthquakes are quantified according to the amount of energy they produce (examples are 2.1, 5.0, 6.7). What type of data is that?

6. 地震按其所释放的能量大小量化(例如 2.1、5.0、6.7)。那是何种类型的数据?

7. What contributed to the large number of deaths in 2010? In 2004? Explain.

7. 是什么导致了 2010 年(以及 2004 年)的大量死亡?请解释。

*For the following four exercises, determine the type of sampling used (simple random, stratified, systematic, cluster, or convenience).*

(在接下来的四道练习中,判断所使用的抽样类型(简单随机、分层、系统、整群或方便)。)

12\.

12.

A group of test subjects is divided into twelve groups; then four of the groups are chosen at random.

一组受试对象被划分为十二个小组,然后随机抽取其中四个小组。

13.

13.

A market researcher polls every tenth person who walks into a store.

一位市场调研人员对走进商店的每第十个人进行访谈。

14\.

14.

The first 50 people who walk into a sporting event are polled on their television preferences.

对走进一场体育赛事的前 50 个人就其电视偏好进行访谈。

15.

15.

A computer generates 100 random numbers, and 100 people whose names correspond with the numbers on the list are chosen.

计算机生成 100 个随机数,然后选取名单上姓名与这些数字对应的 100 个人。

*Use the following information to answer the next seven exercises:* Studies are often done by pharmaceutical companies to determine the effectiveness of a treatment program. Suppose that a new AIDS antibody drug is currently under study. It is given to patients once the AIDS symptoms have revealed themselves. Of interest is the average (mean) length of time in months patients live once starting the treatment. Two researchers each follow a different set of 40 AIDS patients from the start of treatment until their deaths. The following data (in months) are collected.

(用以下信息回答接下来的七道练习。)制药公司常开展研究以确定某个治疗方案的疗效。假设目前一种新型艾滋病抗体药物正在研究中。患者在艾滋病症状显现后开始用药。研究者关注的是患者开始治疗后存活时间(以月计)的平均数。两位研究者各自追踪一组不同的 40 名艾滋病患者,从治疗开始直至其死亡。收集到如下数据(单位:月)。

Researcher A: 3; 4; 11; 15; 16; 17; 22; 44; 37; 16; 14; 24; 25; 15; 26; 27; 33; 29; 35; 44; 13; 21; 22; 10; 12; 8; 40; 32; 26; 27; 31; 34; 29; 17; 8; 24; 18; 47; 33; 34

研究者 A:3; 4; 11; 15; 16; 17; 22; 44; 37; 16; 14; 24; 25; 15; 26; 27; 33; 29; 35; 44; 13; 21; 22; 10; 12; 8; 40; 32; 26; 27; 31; 34; 29; 17; 8; 24; 18; 47; 33; 34

Researcher B: 3; 14; 11; 5; 16; 17; 28; 41; 31; 18; 14; 14; 26; 25; 21; 22; 31; 2; 35; 44; 23; 21; 21; 16; 12; 18; 41; 22; 16; 25; 33; 34; 29; 13; 18; 24; 23; 42; 33; 29

研究者 B:3; 14; 11; 5; 16; 17; 28; 41; 31; 18; 14; 14; 26; 25; 21; 22; 31; 2; 35; 44; 23; 21; 21; 16; 12; 18; 41; 22; 16; 25; 33; 34; 29; 13; 18; 24; 23; 42; 33; 29

16\.

16.

Complete the tables using the data provided:

利用所提供的数据补全表格:
Table 1.27 Researcher A
Survival Length (in months)FrequencyRelative FrequencyCumulative Relative Frequency
0.5–6.5
6.5–12.5
12.5–18.5
18.5–24.5
24.5–30.5
30.5–36.5
36.5–42.5
42.5–48.5
表 1.27 研究者 A
存活时长(月)频数相对频数累计相对频数
0.5–6.5
6.5–12.5
12.5–18.5
18.5–24.5
24.5–30.5
30.5–36.5
36.5–42.5
42.5–48.5
Table 1.28 Researcher B
Survival Length (in months)FrequencyRelative FrequencyCumulative Relative Frequency
0.5–6.5
6.5–12.5
12.5–18.5
18.5–24.5
24.5–30.5
30.5–36.5
36.5-45.5
表 1.28 研究者 B
存活时长(月)频数相对频数累计相对频数
0.5–6.5
6.5–12.5
12.5–18.5
18.5–24.5
24.5–30.5
30.5–36.5
36.5-45.5

17.

17.

Determine what the key term data refers to in the above example for Researcher A.

确定在上述研究者 A 的例子中,术语“数据”所指的对象。

18\.

18.

List two reasons why the data may differ.

列出数据可能不同的两条原因。

19.

19.

Can you tell if one researcher is correct and the other one is incorrect? Why?

你能判断哪位研究者正确、另一位错误吗?为什么?

20\.

20.

Would you expect the data to be identical? Why or why not?

你预期数据会完全相同吗?为什么相同或不相同?

21.

21.

Suggest at least two methods the researchers might use to gather random data.

提出研究者可能用以收集随机数据的至少两种方法。

22\.

22.

Suppose that the first researcher conducted his survey by randomly choosing one state in the nation and then randomly picking 40 patients from that state. What sampling method would that researcher have used?

假设第一位研究者通过随机选取全国某一个州,再从该州随机抽选 40 名患者来进行调查。该研究者使用的是何种抽样方法?

23.

23.

Suppose that the second researcher conducted his survey by choosing 40 patients he knew. What sampling method would that researcher have used? What concerns would you have about this data set, based upon the data collection method?

假设第二位研究者通过选取他认识的 40 名患者进行调查。该研究者使用的是何种抽样方法?基于这种数据收集方式,你对该数据集有何顾虑?

*Use the following data to answer the next five exercises:* Two researchers are gathering data on hours of video games played by school-aged children and young adults. They each randomly sample different groups of 150 students from the same school. They collect the following data.

(用以下数据回答接下来的五道练习。)两位研究者正在收集学龄儿童与年轻成年人玩电子游戏时长的数据。他们各自从同一所学校随机抽取不同的 150 名学生组成的群体。收集到如下数据。
Table 1.29 Researcher A
Hours Played per WeekFrequencyRelative FrequencyCumulative Relative Frequency
0–2260.170.17
2–4300.200.37
4–6490.330.70
6–8250.170.87
8–10120.080.95
10–1280.051
表 1.29 研究者 A
每周游戏时长(小时)频数相对频数累计相对频数
0–2260.170.17
2–4300.200.37
4–6490.330.70
6–8250.170.87
8–10120.080.95
10–1280.051
Table 1.30 Researcher B
Hours Played per WeekFrequencyRelative FrequencyCumulative Relative Frequency
0–2480.320.32
2–4510.340.66
4–6240.160.82
6–8120.080.90
8–10110.070.97
10–1240.031
表 1.30 研究者 B
每周游戏时长(小时)频数相对频数累计相对频数
0–2480.320.32
2–4510.340.66
4–6240.160.82
6–8120.080.90
8–10110.070.97
10–1240.031

24.

24.

Give a reason why the data may differ.

给出数据可能不同的一个原因。

25.

25.

Would the sample size be large enough if the population is the students in the school?

若总体为该学校的学生,样本量是否足够大?

26\.

26.

Would the sample size be large enough if the population is school-aged children and young adults in the United States?

若总体为美国学龄儿童与年轻成年人,样本量是否足够大?

27.

27.

Researcher A concludes that most students play video games between four and six hours each week. Researcher B concludes that most students play video games between two and four hours each week. Who is correct?

研究者 A 得出大多数学生每周玩电子游戏 4 至 6 小时,研究者 B 得出大多数学生每周玩 2 至 4 小时。谁是正确的?

28\.

28.

As part of a way to reward students for participating in the survey, the researchers gave each student a gift card to a video game store. Would this affect the data if students knew about the award before the study?

作为奖励学生参与调查的一种方式,研究者向每名学生赠送了一张电子游戏商店的礼品卡。如果学生在研究前就知道了这项奖励,这会影响数据吗?

*Use the following data to answer the next five exercises:* A pair of studies was performed to measure the effectiveness of a new software program designed to help stroke patients regain their problem-solving skills. Patients were asked to use the software program twice a day, once in the morning and once in the evening. The studies observed 200 stroke patients recovering over a period of several weeks. The first study collected the data in Table 1.31. The second study collected the data in Table 1.32.

(用以下数据回答接下来的五道练习。)一对研究被用于评估一款帮助中风患者恢复解决问题能力的新软件程序的效果。要求患者每天使用该软件两次,一次在早晨、一次在傍晚。两项研究在数周时间内观察了 200 名康复中的中风患者。第一项研究的数据收集于表 1.31,第二项研究的数据收集于表 1.32。
Table 1.31
GroupShowed improvementNo improvementDeterioration
Used program1424315
Did not use program7211018
表 1.31
组别有改善无改善恶化
使用程序1424315
未使用程序7211018
Table 1.32
GroupShowed improvementNo improvementDeterioration
Used program1057419
Did not use program899912
表 1.32
组别有改善无改善恶化
使用程序1057419
未使用程序899912

29.

29.

Given what you know, which study is correct?

根据你的了解,哪项研究是正确的?

30\.

30.

The first study was performed by the company that designed the software program. The second study was performed by the American Medical Association. Which study is more reliable?

第一项研究由设计该软件程序的公司实施。第二项研究由美国医学协会实施。哪项研究更可靠?

31.

31.

Both groups that performed the study concluded that the software works. Is this accurate?

实施研究的两个小组都得出结论称该软件有效。这准确吗?

32\.

32.

The company takes the two studies as proof that their software causes mental improvement in stroke patients. Is this a fair statement?

该公司将这两项研究作为其软件能改善中风患者心智能力的证明。这种说法公平吗?

33.

33.

Patients who used the software were also a part of an exercise program whereas patients who did not use the software were not. Does this change the validity of the conclusions from Exercise 1.31?

使用软件的患者同时也参与了一项锻炼计划,而未使用软件的患者则没有。这会改变练习 1.31 中结论的有效性吗?

34\.

34.

Is a sample size of 1,000 a reliable measure for a population of 5,000?

对 5,000 的总体而言,1,000 的样本量是否可靠?

35.

35.

Is a sample of 500 volunteers a reliable measure for a population of 2,500?

对 2,500 的总体而言,500 名志愿者组成的样本是否可靠?

36\.

36.

A question on a survey reads: "Do you prefer the delicious taste of Brand X or the taste of Brand Y?" Is this a fair question?

调查问卷上有一道题:“你更喜欢 X 品牌的美味还是 Y 品牌的味道?”这是一个公平的问题吗?

37.

37.

Is a sample size of two representative of a population of five?

对 5 的总体而言,样本量为 2 是否具有代表性?

38\.

38.

Is it possible for two experiments to be well run with similar sample sizes to get different data?

两项设计良好、样本量相近的实验是否可能得到不同的数据?

1.3 Frequency, Frequency Tables, and Levels of Measurement 1.3 频数、频数表与测量尺度

39.

39.

What type of measure scale is being used? Nominal, ordinal, interval or ratio.

使用的是哪种测量尺度?名义、顺序、区间还是比率。

1. High school soccer players classified by their athletic ability: Superior, Average, Above average

1. 按运动能力分类的高中足球运动员:优秀、一般、高于一般

2. Baking temperatures for various main dishes: 350, 400, 325, 250, 300

2. 各种主菜的烘焙温度:350、400、325、250、300

3. The colors of crayons in a 24-crayon box

3. 一盒24色蜡笔的颜色

4. Social security numbers

4. 社会保障号码

5. Incomes measured in dollars

5. 以美元计量的收入

6. A satisfaction survey of a social website by number: 1 = very satisfied, 2 = somewhat satisfied, 3 = not satisfied

6. 某社交网站的满意度调查(用数字表示):1 = 非常满意,2 = 比较满意,3 = 不满意

7. Political outlook: extreme left, left-of-center, right-of-center, extreme right

7. 政治倾向:极左、中左、中右、极右

8. Time of day on an analog watch

8. 指针式手表上显示的时间

9. The distance in miles to the closest grocery store

9. 到最近杂货店的英里距离

10. The dates 1066, 1492, 1644, 1947, and 1944

10. 日期 1066、1492、1644、1947 和 1944

11. The heights of 21–65 year-old women

11. 21–65 岁女性的身高

12. Common letter grades: A, B, C, D, and F

12. 常用字母成绩:A、B、C、D 和 F

1.4 Experimental Design and Ethics 1.4 实验设计与伦理

40\.

40.

Design an experiment. Identify the explanatory and response variables. Describe the population being studied and the experimental units. Explain the treatments that will be used and how they will be assigned to the experimental units. Describe how blinding and placebos may be used to counter the power of suggestion.

设计一个实验。明确解释变量与响应变量。描述所研究的总体与实验单元。说明将采用的处理及其如何分配给实验单元。描述如何使用盲法与安慰剂来抵消暗示的力量。

41.

41.

Discuss potential violations of the rule requiring informed consent.

讨论对"要求知情同意"这一规则的可能违反情形。

1. Inmates in a correctional facility are offered good behavior credit in return for participation in a study.

1. 惩教机构中的在押人员若参与某项研究,可获得良好行为积分作为回报。

2. A research study is designed to investigate a new children’s allergy medication.

2. 一项研究旨在调查一种新的儿童过敏药物。

3. Participants in a study are told that the new medication being tested is highly promising, but they are not told that only a small portion of participants will receive the new medication. Others will receive placebo treatments and traditional treatments.

3. 研究参与者被告知所测试的新药很有前景,但未被告知只有一小部分参与者会接受新药,其余人接受安慰剂治疗和传统治疗。

Homework 作业

1.1 Definitions of Statistics, Probability, and Key Terms 1.1 统计学、概率与关键术语的定义

*For each of the following eight exercises, identify: a. the population, b. the sample, c. the parameter, d. the statistic, e. the variable, and f. the data. Give examples where appropriate.*

*对于以下八个练习,请识别:a. 总体,b. 样本,c. 参数,d. 统计量,e. 变量,f. 数据。酌情给出示例。*

42\.

42.

A fitness center is interested in the mean amount of time a client exercises in the center each week.

某健身中心关注其会员每周在该中心锻炼的平均时长。

43.

43.

Ski resorts are interested in the mean age that children take their first ski and snowboard lessons. They need this information to plan their ski classes optimally.

滑雪场关注儿童首次参加滑雪与单板滑雪课程的平均年龄。它们需要这些信息来优化课程安排。

44\.

44.

A cardiologist is interested in the mean recovery period of her patients who have had heart attacks.

一位心脏病专家关注其心脏病发作患者的平均康复期。

45.

45.

Insurance companies are interested in the mean health costs each year of their clients, so that they can determine the costs of health insurance.

保险公司关注其客户每年的平均医疗成本,以便确定健康保险的费用。

46\.

46.

A politician is interested in the proportion of voters in his district who think he is doing a good job.

一位政客关注其选区内认为他工作出色的选民比例。

47.

47.

A marriage counselor is interested in the proportion of clients she counsels who stay married.

一位婚姻咨询师关注其所咨询客户中维持婚姻的比例。

48\.

48.

Political pollsters may be interested in the proportion of people who will vote for a particular cause.

政治民意调查者可能关注将为某一特定主张投票的人数比例。

49.

49.

A marketing company is interested in the proportion of people who will buy a particular product.

一家营销公司关注将购买某一特定产品的人数比例。

*Use the following information to answer the next three exercises:* A Lake Tahoe Community College instructor is interested in the mean number of days Lake Tahoe Community College math students are absent from class during a quarter.

*利用以下信息回答接下来三道练习:*一位塔霍湖社区学院的教师关注该学院数学专业学生在一个学期中缺课的平均天数。

50\.

50.

What is the population she is interested in?

她所关注的总体是什么?

1. all Lake Tahoe Community College students

1. 塔霍湖社区学院全体学生

2. all Lake Tahoe Community College English students

2. 塔霍湖社区学院全体英语专业学生

3. all Lake Tahoe Community College students in her classes

3. 她所授课班级中的塔霍湖社区学院全体学生

4. all Lake Tahoe Community College math students

4. 塔霍湖社区学院全体数学专业学生

51.

51.

Consider the following:

考虑以下:

$X$ = number of days a Lake Tahoe Community College math student is absent

$X$ = 塔霍湖社区学院数学专业学生缺课的天数

In this case, *X* is an example of a:

在此情形下,*X* 是下列哪项的例子:

1. variable.

1. 变量。

2. population.

2. 总体。

3. statistic.

3. 统计量。

4. data.

4. 数据。

52\.

52.

The instructor’s sample produces a mean number of days absent of 3.5 days. This value is an example of a:

该教师的样本得出缺课平均天数为 3.5 天。这个数值是下列哪项的例子:

1. parameter.

1. 参数。

2. data.

2. 数据。

3. statistic.

3. 统计量。

4. variable.

4. 变量。

1.2 Data, Sampling, and Variation in Data and Sampling 1.2 数据、抽样与数据及抽样中的变异

*For the following exercises, identify the type of data that would be used to describe a response (quantitative discrete, quantitative continuous, or qualitative), and give an example of the data.*

*对于以下练习,请识别用于描述某一回答的数据类型(离散定量、连续定量或定性),并给出该数据的一个示例。*

53.

53.

number of tickets sold to a concert

一场音乐会售出的门票数量

54\.

54.

percent of body fat

体脂百分比

55.

55.

favorite baseball team

最喜欢的棒球队

56\.

56.

time in line to buy groceries

排队购买食品杂货的时间

57.

57.

number of students enrolled at Evergreen Valley College

常青谷学院注册学生人数

58\.

58.

most-watched television show

收视率最高的电视节目

59.

59.

brand of toothpaste

牙膏品牌

60\.

60.

distance to the closest movie theatre

到最近电影院的距离

61.

61.

age of executives in Fortune 500 companies

财富500强企业高管年龄

62\.

62.

number of competing computer spreadsheet software packages

相互竞争的计算机电子表格软件包数量

*Use the following information to answer the next two exercises:* A study was done to determine the age, number of times per week, and the duration (amount of time) of resident use of a local park in San Jose. The first house in the neighborhood around the park was selected randomly and then every 8th house in the neighborhood around the park was interviewed.

*利用以下信息回答接下来两道练习:*一项研究旨在确定圣何塞当地一个公园周边居民的使用情况,包括年龄、每周使用次数以及使用时长(时间量)。公园周边社区的第一户被随机选中,之后该社区每隔第8户接受访谈。

63.

63.

“Number of times per week” is what type of data?

"每周使用次数"是什么类型的数据?

1. qualitative

1. 定性

2. quantitative discrete

2. 离散定量

3. quantitative continuous

3. 连续定量

64\.

64.

“Duration (amount of time)” is what type of data?

"时长(时间量)"是什么类型的数据?

1. qualitative

1. 定性

2. quantitative discrete

2. 离散定量

3. quantitative continuous

3. 连续定量

65.

65.

Airline companies are interested in the consistency of the number of babies on each flight, so that they have adequate safety equipment. Suppose an airline conducts a survey. Over Thanksgiving weekend, it surveys six flights from Boston to Salt Lake City to determine the number of babies on the flights. It determines the amount of safety equipment needed by the result of that study.

航空公司关注每趟航班上婴儿数量的一致性,以便配备充足的安全设备。假设某航空公司开展一项调查。在感恩节周末,它调查了从波士顿到盐湖城的六趟航班,以确定航班上的婴儿数量。它根据该研究结果来确定所需安全设备的数量。

1. Using complete sentences, list three things wrong with the way the survey was conducted.

1. 用完整的句子列出该调查方式存在的三个问题。

2. Using complete sentences, list three ways that you would improve the survey if it were to be repeated.

2. 用完整的句子列出若该调查重复进行,你会改进它的三种方式。

66\.

66.

Suppose you want to determine the mean number of students per statistics class in your state. Describe a possible sampling method in three to five complete sentences. Make the description detailed.

假设你想确定你所在州每个统计班的平均学生数。请用三到五个完整的句子描述一种可能的抽样方法,并尽量详细。

67.

67.

Suppose you want to determine the mean number of cans of soda drunk each month by students in their twenties at your school. Describe a possible sampling method in three to five complete sentences. Make the description detailed.

假设你想确定你学校二十多岁学生每月饮用的汽水罐数平均值。请用三到五个完整的句子描述一种可能的抽样方法,并尽量详细。

68\.

68.

List some practical difficulties involved in getting accurate results from a telephone survey.

列出电话调查中获取准确结果所面临的一些实际困难。

69.

69.

List some practical difficulties involved in getting accurate results from a mailed survey.

列出邮寄调查中获取准确结果所面临的一些实际困难。

70\.

70.

With your classmates, brainstorm some ways you could overcome these problems if you needed to conduct a phone or mail survey.

与同学一起头脑风暴:如果你需要进行电话或邮寄调查,可以想出哪些办法克服这些问题。

71.

71.

The instructor takes her sample by gathering data on five randomly selected students from each Lake Tahoe Community College math class. The type of sampling she used is

该教师通过从塔霍湖社区学院每个数学班随机选取五名学生收集数据来抽取样本。她使用的抽样类型是

1. cluster sampling

1. 整群抽样

2. stratified sampling

2. 分层抽样

3. simple random sampling

3. 简单随机抽样

4. convenience sampling

4. 方便抽样

72\.

72.

A study was done to determine the age, number of times per week, and the duration (amount of time) of residents using a local park in San Jose. The first house in the neighborhood around the park was selected randomly and then every eighth house in the neighborhood around the park was interviewed. The sampling method was:

一项研究旨在确定圣何塞当地一个公园使用者的年龄、每周使用次数以及使用时长(时间量)。公园周边社区的第一户被随机选中,之后该社区每隔第8户接受访谈。抽样方法是:

1. simple random

1. 简单随机

2. systematic

2. 系统

3. stratified

3. 分层

4. cluster

4. 整群

73.

73.

Name the sampling method used in each of the following situations:

指出下列各种情形中使用的抽样方法:

1. A woman in the airport is handing out questionnaires to travelers asking them to evaluate the airport’s service. She does not ask travelers who are hurrying through the airport with their hands full of luggage, but instead asks all travelers who are sitting near gates and not taking naps while they wait.

1. 机场中一名女子向旅客发放问卷,请他们评价机场服务。她不询问那些手提满行李匆匆穿行机场的旅客,而是询问所有坐在登机口附近、等待时未打盹的旅客。

2. A teacher wants to know if her students are doing homework, so she randomly selects rows two and five and then calls on all students in row two and all students in row five to present the solutions to homework problems to the class.

2. 一位教师想知道学生是否在做作业,于是她随机选取第二排和第五排,然后请第二排和第五排的所有学生向全班展示作业题的解答。

3. The marketing manager for an electronics chain store wants information about the ages of its customers. Over the next two weeks, at each store location, 100 randomly selected customers are given questionnaires to fill out asking for information about age, as well as about other variables of interest.

3. 一家电子产品连锁店的市场经理想了解顾客年龄。在接下来两周内,每个门店地点都会向100名随机选取的顾客发放问卷,询问年龄及其他相关变量。

4. The librarian at a public library wants to determine what proportion of the library users are children. The librarian has a tally sheet on which she marks whether books are checked out by an adult or a child. She records this data for every fourth patron who checks out books.

4. 一家公共图书馆的馆员想确定图书馆使用者中儿童的比例。她有一张记录表,标记借出的书是由成人还是儿童借出。她对每个第四位借书的读者记录这一数据。

5. A political party wants to know the reaction of voters to a debate between the candidates. The day after the debate, the party’s polling staff calls 1,200 randomly selected phone numbers. If a registered voter answers the phone or is available to come to the phone, that registered voter is asked whom he or she intends to vote for and whether the debate changed his or her opinion of the candidates.

5. 一个政党想了解选民对候选人辩论的反应。辩论次日,该党民调人员拨打1,200个随机选取的电话号码。如果有登记选民接听电话或可以接听,就询问其打算投给谁,以及辩论是否改变了其对候选人的看法。

74\.

74.

A “random survey” was conducted of 3,274 people of the “microprocessor generation” (people born since 1971, the year the microprocessor was invented). It was reported that 48% of those individuals surveyed stated that if they had \$2,000 to spend, they would use it for computer equipment. Also, 66% of those surveyed considered themselves relatively savvy computer users.

一项"随机调查"对3,274名"微处理器一代"(1971年即微处理器发明之年出生的人)展开。据报道,受访人群中48%的人表示,如果有 2,000 美元可花,他们会将其用于计算机设备。此外,66%的受访者认为自己是相当精通计算机的用户。

1. Do you consider the sample size large enough for a study of this type? Why or why not?

1. 你认为该样本量对这类研究而言足够大吗?为什么或为什么不?

2. Based on your “gut feeling,” do you believe the percents accurately reflect the U.S. population for those individuals born since 1971? If not, do you think the percents of the population are actually higher or lower than the sample statistics? Why?

2. 凭你的"直觉",你认为这些百分比准确反映了1971年以来出生人群的美国总体吗?如果不准确,你认为总体中的百分比实际上高于还是低于样本统计量?为什么?

Additional information: The survey, reported by Intel Corporation, was filled out by individuals who visited the Los Angeles Convention Center to see the Smithsonian Institute's road show called “America’s Smithsonian.”

补充信息:该调查由英特尔公司报告,填写者为参观洛杉矶会议中心史密森学会路演"美国的史密森"的人士。

3. With this additional information, do you feel that all demographic and ethnic groups were equally represented at the event? Why or why not?

3. 有了这一补充信息,你认为所有人口统计与族群在该活动中都得到了平等代表吗?为什么或为什么不?

4. With the additional information, comment on how accurately you think the sample statistics reflect the population parameters.

4. 有了这一补充信息,请评论你认为样本统计量在多大程度上准确反映了总体参数。

75.

75.

The Well-Being Index is a survey that follows trends of U.S. residents on a regular basis. There are six areas of health and wellness covered in the survey: Life Evaluation, Emotional Health, Physical Health, Healthy Behavior, Work Environment, and Basic Access. Some of the questions used to measure the Index are listed below.

幸福指数是一项定期追踪美国居民趋势的调查。该调查涵盖健康与保健的六个领域:生活评价、情绪健康、身体健康、健康行为、工作环境和基本保障。下面列出用于衡量该指数的一些问题。

Identify the type of data obtained from each question used in this survey: qualitative, quantitative discrete, or quantitative continuous.

指出该调查中每个问题所获取数据的数据类型:定性、离散定量或连续定量。

1. Do you have any health problems that prevent you from doing any of the things people your age can normally do?

1. 你是否有任何健康问题,使你无法做同龄人通常能做的一些事情?

2. During the past 30 days, for about how many days did poor health keep you from doing your usual activities?

2. 在过去30天里,大约有多少天因健康不佳使你无法从事日常活动?

3. In the last seven days, on how many days did you exercise for 30 minutes or more?

3. 在过去七天里,有多少天你进行了30分钟或以上的锻炼?

4. Do you have health insurance coverage?

4. 你是否有健康保险?

76\.

76.

In advance of the 1936 Presidential Election, a magazine titled Literary Digest released the results of an opinion poll predicting that the republican candidate Alf Landon would win by a large margin. The magazine sent post cards to approximately 10,000,000 prospective voters. These prospective voters were selected from the subscription list of the magazine, from automobile registration lists, from phone lists, and from club membership lists. Approximately 2,300,000 people returned the postcards.

在1936年总统大选前,一本名为《文学文摘》的杂志发布了民意调查结果,预测共和党候选人阿尔夫·兰登将以大幅优势获胜。该杂志向约10,000,000名潜在选民寄送明信片。这些潜在选民选自该杂志的订阅名单、汽车登记名单、电话名单和俱乐部会员名单。约2,300,000人寄回了明信片。

1. Think about the state of the United States in 1936. Explain why a sample chosen from magazine subscription lists, automobile registration lists, phone books, and club membership lists was not representative of the population of the United States at that time.

1. 思考1936年美国的状况。解释为何选自杂志订阅名单、汽车登记名单、电话簿和俱乐部会员名单的样本不能代表当时美国的总体。

2. What effect does the low response rate have on the reliability of the sample?

2. 低回答率对样本可靠性有何影响?

3. Are these problems examples of sampling error or nonsampling error?

3. 这些问题是抽样误差还是非抽样误差的例子?

4. During the same year, George Gallup conducted his own poll of 30,000 prospective voters. These researchers used a method they called "quota sampling" to obtain survey answers from specific subsets of the population. Quota sampling is an example of which sampling method described in this module?

4. 同年,乔治·盖洛普对30,000名潜在选民开展了自己的民调。这些研究者使用了一种称为"配额抽样"的方法,从总体的特定子集中获取调查答案。配额抽样是本模块描述的哪种抽样方法的例子?

77.

77.

Crime-related and demographic statistics for 47 US states in 1960 were collected from government agencies, including the FBI's *Uniform Crime Report*. One analysis of this data found a strong connection between education and crime indicating that higher levels of education in a community correspond to higher crime rates.

1960年美国47个州的犯罪相关与人口统计数据来自政府机构,包括联邦调查局的*《统一犯罪报告》*。对该数据的一项分析发现教育与犯罪之间存在强关联,表明社区教育水平越高,犯罪率越高。

Which of the potential problems with samples discussed in 1.2 Data, Sampling, and Variation in Data and Sampling could explain this connection?

在1.2"数据、抽样与数据及抽样中的变异"中讨论的样本潜在问题中,哪一个可以解释这种关联?

78\.

78.

YouPolls is a website that allows anyone to create and respond to polls. One question posted April 15 asks:

YouPolls 是一个允许任何人创建并回答投票的网站。4月15日发布的一个问题问道:

“Do you feel happy paying your taxes when members of the Obama administration are allowed to ignore their tax liabilities?” (lastbaldeagle. 2013. On Tax Day, House to Call for Firing Federal Workers Who Owe Back Taxes. Opinion poll posted online at: http://www.youpolls.com/details.aspx?id=12328 (accessed May 1, 2013).)

"当奥巴马政府成员被允许忽视其纳税义务时,你是否为纳税感到高兴?"(lastbaldeagle. 2013. 在纳税日,众议院呼吁解雇欠税的联邦工作人员。在线发布的民意调查见:http://www.youpolls.com/details.aspx?id=12328(访问于2013年5月1日)。)

As of April 25, 11 people responded to this question. Each participant answered “NO!”

截至4月25日,有11人回答了该问题。每位参与者都回答"不!"

Which of the potential problems with samples discussed in this module could explain this connection?

在本模块讨论的样本潜在问题中,哪一个可以解释这种关联?

79.

79.

A scholarly article about response rates begins with the following quote:

一篇关于回答率的学术文章以以下引文开头:

“Declining contact and cooperation rates in random digit dial (RDD) national telephone surveys raise serious concerns about the validity of estimates drawn from such research.”(Scott Keeter et al., “Gauging the Impact of Growing Nonresponse on Estimates from a National RDD Telephone Survey,” Public Opinion Quarterly 70 no. 5 (2006), http://poq.oxfordjournals.org/content/70/5/759.full (accessed May 1, 2013).)

"随机数字拨号(RDD)全国电话调查中联系率与合作率的下降,引发了对由此类研究得出的估计有效性的严重关切。"(Scott Keeter 等,"衡量日益增长的无回答对全国RDD电话调查估计的影响",《公共舆论季刊》70卷5期(2006年),http://poq.oxfordjournals.org/content/70/5/759.full(访问于2013年5月1日)。)

The Pew Research Center for People and the Press admits:

皮尤研究中心(人民与新闻界)承认:

“The percentage of people we interview – out of all we try to interview – has been declining over the past decade or more.” (Frequently Asked Questions, Pew Research Center for the People & the Press, http://www.people-press.org/methodology/frequently-asked-questions/#dont-you-have-trouble-getting-people-to-answer-your-polls (accessed May 1, 2013).)

"我们访谈的人数——占我们尝试访谈总人数的比例——在过去十年或更久以来一直在下降。"(常见问题,皮尤研究中心(人民与新闻界),http://www.people-press.org/methodology/frequently-asked-questions/#dont-you-have-trouble-getting-people-to-answer-your-polls(访问于2013年5月1日)。)

1. What are some reasons for the decline in response rate over the past decade?

1. 过去十年回答率下降的原因有哪些?

2. Explain why researchers are concerned with the impact of the declining response rate on public opinion polls.

2. 解释研究者为何关注回答率下降对民意调查的影响。

1.3 Frequency, Frequency Tables, and Levels of Measurement 1.3 频数、频数表与测量尺度

80\.

80.

Fifty part-time students were asked how many courses they were taking this term. The (incomplete) results are shown below:

五十名非全日制学生被问及他们本学期修读了多少门课程。以下显示的是(不完整的)结果:

| \# of Courses | Frequency | Relative Frequency | Cumulative Relative Frequency |

| # 课程数 | 频数 | 相对频数 | 累计相对频数 |

|---------------|-----------|--------------------|-------------------------------|

|---------------|-----------|--------------------|-------------------------------|

| 1 | 30 | 0.6 | |

| 1 | 30 | 0.6 | |

| 2 | 15 | | |

| 2 | 15 | | |

| 3 | | | |

| 3 | | | |

Table 1.33 Part-time Student Course Loads

表 1.33 非全日制学生课程负担

1. Fill in the blanks in Table 1.33.

1. 填写表 1.33 中的空白。

2. What percent of students take exactly two courses?

2. 修读恰好两门课程的学生占百分之几?

3. What percent of students take one or two courses?

3. 修读一门或两门课程的学生占百分之几?

81\.

81.

Sixty adults with gum disease were asked the number of times per week they used to floss before their diagnosis. The (incomplete) results are shown in Table 1.34.

六十名患有牙龈炎的成年人被问及在被诊断前他们每周使用牙线的次数。以下(不完整的)结果显示于表 1.34。

| \# Flossing per Week | Frequency | Relative Frequency | Cumulative Relative Freq. |

| # 每周牙线使用次数 | 频数 | 相对频数 | 累计相对频数 |

|----------------------|-----------|--------------------|---------------------------|

|----------------------|-----------|--------------------|---------------------------|

| 0 | 27 | 0.4500 | |

| 0 | 27 | 0.4500 | |

| 1 | 18 | | |

| 1 | 18 | | |

| 3 | | | 0.9333 |

| 3 | | | 0.9333 |

| 6 | 3 | 0.0500 | |

| 6 | 3 | 0.0500 | |

| 7 | 1 | 0.0167 | |

| 7 | 1 | 0.0167 | |

Table 1.34 Flossing Frequency for Adults with Gum Disease

表 1.34 患牙龈炎成年人的牙线使用频数

1. Fill in the blanks in Table 1.34.

1. 填写表 1.34 中的空白。

2. What percent of adults flossed six times per week?

2. 每周使用牙线六次的成年人占百分之几?

3. What percent flossed at most three times per week?

3. 每周最多使用牙线三次的人占百分之几?

82\.

82.

Nineteen immigrants to the U.S were asked how many years, to the nearest year, they have lived in the U.S. The data are as follows: 2; 5; 7; 2; 2; 10; 20; 15; 0; 7; 0; 20; 5; 12; 15; 12; 4; 5; 10 .

十九名移民到美国的人被问及他们已在美国居住了多少年(精确到年)。数据如下:2;5;7;2;2;10;20;15;0;7;0;20;5;12;15;12;4;5;10。

Table 1.35 was produced.

据此生成了表 1.35。

| Data | Frequency | Relative Frequency | Cumulative Relative Frequency |

| 数据 | 频数 | 相对频数 | 累计相对频数 |

|------|-----------|--------------------|-------------------------------|

|------|-----------|--------------------|-------------------------------|

| 0 | 2 | $\frac{2}{19}$ | 0.1053 |

| 0 | 2 | $\frac{2}{19}$ | 0.1053 |

| 2 | 3 | $\frac{3}{19}$ | 0.2632 |

| 2 | 3 | $\frac{3}{19}$ | 0.2632 |

| 4 | 1 | $\frac{1}{19}$ | 0.3158 |

| 4 | 1 | $\frac{1}{19}$ | 0.3158 |

| 5 | 3 | $\frac{3}{19}$ | 0.4737 |

| 5 | 3 | $\frac{3}{19}$ | 0.4737 |

| 7 | 2 | $\frac{2}{19}$ | 0.5789 |

| 7 | 2 | $\frac{2}{19}$ | 0.5789 |

| 10 | 2 | $\frac{2}{19}$ | 0.6842 |

| 10 | 2 | $\frac{2}{19}$ | 0.6842 |

| 12 | 2 | $\frac{2}{19}$ | 0.7895 |

| 12 | 2 | $\frac{2}{19}$ | 0.7895 |

| 15 | 1 | $\frac{1}{19}$ | 0.8421 |

| 15 | 1 | $\frac{1}{19}$ | 0.8421 |

| 20 | 1 | $\frac{1}{19}$ | 1.0000 |

| 20 | 1 | $\frac{1}{19}$ | 1.0000 |

Table 1.35 Frequency of Immigrant Survey Responses

表 1.35 移民调查回答的频数

1. Fix the errors in Table 1.35. Also, explain how someone might have arrived at the incorrect number(s).

1. 修正表 1.35 中的错误。并解释有人是如何得出这些错误数字的。

2. Explain what is wrong with this statement: “47 percent of the people surveyed have lived in the U.S. for 5 years.”

2. 解释这句话的错误之处:“47% 的受访者在美居住了 5 年。”

3. Fix the statement in b to make it correct.

3. 修正上述陈述 (b) 使其正确。

4. What fraction of the people surveyed have lived in the U.S. five or seven years?

4. 受访者在美居住了 5 年或 7 年的占几分之几?

5. What fraction of the people surveyed have lived in the U.S. at most 12 years?

5. 受访者在美居住至多 12 年的占几分之几?

6. What fraction of the people surveyed have lived in the U.S. fewer than 12 years?

6. 受访者在美居住少于 12 年的占几分之几?

7. What fraction of the people surveyed have lived in the U.S. from five to 20 years, inclusive?

7. 受访者在美居住 5 至 20 年(含)的占几分之几?

83\.

83.

How much time does it take to travel to work? Table 1.36 shows the mean commute time by state for workers at least 16 years old who are not working at home. Find the mean travel time, and round off the answer properly.

通勤上班需要多长时间?表 1.36 显示了至少 16 岁、且不在家中工作的劳动者各州的平均通勤时间。求出平均通勤时间,并正确地对答案进行四舍五入。

| | | | | | | | | | |

| | | | | | | | | | |

|------|------|------|------|------|------|------|------|------|------|

|------|------|------|------|------|------|------|------|------|------|

| 24.0 | 24.3 | 25.9 | 18.9 | 27.5 | 17.9 | 21.8 | 20.9 | 16.7 | 27.3 |

| 24.0 | 24.3 | 25.9 | 18.9 | 27.5 | 17.9 | 21.8 | 20.9 | 16.7 | 27.3 |

| 18.2 | 24.7 | 20.0 | 22.6 | 23.9 | 18.0 | 31.4 | 22.3 | 24.0 | 25.5 |

| 18.2 | 24.7 | 20.0 | 22.6 | 23.9 | 18.0 | 31.4 | 22.3 | 24.0 | 25.5 |

| 24.7 | 24.6 | 28.1 | 24.9 | 22.6 | 23.6 | 23.4 | 25.7 | 24.8 | 25.5 |

| 24.7 | 24.6 | 28.1 | 24.9 | 22.6 | 23.6 | 23.4 | 25.7 | 24.8 | 25.5 |

| 21.2 | 25.7 | 23.1 | 23.0 | 23.9 | 26.0 | 16.3 | 23.1 | 21.4 | 21.5 |

| 21.2 | 25.7 | 23.1 | 23.0 | 23.9 | 26.0 | 16.3 | 23.1 | 21.4 | 21.5 |

| 27.0 | 27.0 | 18.6 | 31.7 | 23.3 | 30.1 | 22.9 | 23.3 | 21.7 | 18.6 |

| 27.0 | 27.0 | 18.6 | 31.7 | 23.3 | 30.1 | 22.9 | 23.3 | 21.7 | 18.6 |

Table 1.36 84.

表 1.36 84.

*Forbes* magazine published data on the best small firms in 2012. These were firms which had been publicly traded for at least a year, have a stock price of at least \$5 per share, and have reported annual revenue between \$5 million and \$1 billion. Table 1.37 shows the ages of the chief executive officers for the first 60 ranked firms.

《福布斯》杂志发布了 2012 年最佳小型企业的数据。这些企业至少已公开交易一年,每股股价至少 \$5,且年报营收在 \$500 万至 \$10 亿之间。表 1.37 显示了排名前 60 家企业的首席执行官年龄。

| Age | Frequency | Relative Frequency | Cumulative Relative Frequency |

| 年龄 | 频数 | 相对频数 | 累计相对频数 |

|-------|-----------|--------------------|-------------------------------|

|-------|-----------|--------------------|-------------------------------|

| 40–44 | 3 | | |

| 40–44 | 3 | | |

| 45–49 | 11 | | |

| 45–49 | 11 | | |

| 50–54 | 13 | | |

| 50–54 | 13 | | |

| 55–59 | 16 | | |

| 55–59 | 16 | | |

| 60–64 | 10 | | |

| 60–64 | 10 | | |

| 65–69 | 6 | | |

| 65–69 | 6 | | |

| 70–74 | 1 | | |

| 70–74 | 1 | | |

Table 1.37

表 1.37

1. What is the frequency for CEO ages between 54 and 65?

1. 年龄在 54 至 65 岁之间的首席执行官频数是多少?

2. What percentage of CEOs are 65 years or older?

2. 65 岁及以上的首席执行官占百分之几?

3. What is the relative frequency of ages under 50?

3. 50 岁以下年龄的相对频数是多少?

4. What is the cumulative relative frequency for CEOs younger than 55?

4. 55 岁以下首席执行官的累计相对频数是多少?

5. Which graph shows the relative frequency and which shows the cumulative relative frequency?

5. 哪幅图显示相对频数,哪幅图显示累计相对频数?

*Use the following information to answer the next two exercises:* Table 1.38 contains data on hurricanes that have made direct hits on the U.S. Between 1851 and 2004. A hurricane is given a strength category rating based on the minimum wind speed generated by the storm.

*利用以下信息回答接下来的两道习题:* 表 1.38 包含了 1851 年至 2004 年间直接袭击美国的飓风数据。飓风根据风暴产生的最小风速被划分为强度等级。

| Category | Number of Direct Hits | Relative Frequency | Cumulative Frequency |

| 等级 | 直接袭击次数 | 相对频数 | 累计频数 |

|----------|-----------------------|--------------------|----------------------|

|----------|-----------------------|--------------------|----------------------|

| 1 | 109 | 0.3993 | 0.3993 |

| 1 | 109 | 0.3993 | 0.3993 |

| 2 | 72 | 0.2637 | 0.6630 |

| 2 | 72 | 0.2637 | 0.6630 |

| 3 | 71 | 0.2601 | |

| 3 | 71 | 0.2601 | |

| 4 | 18 | | 0.9890 |

| 4 | 18 | | 0.9890 |

| 5 | 3 | 0.0110 | 1.0000 |

| 5 | 3 | 0.0110 | 1.0000 |

| | Total = 273 | | |

| | 总计 = 273 | | |

Table 1.38 Frequency of Hurricane Direct Hits 85.

表 1.38 飓风直接袭击的频数 85.

What is the relative frequency of direct hits that were category 4 hurricanes?

属于 4 级飓风的直袭相对频数是多少?

1. 0.0768

1. 0.0768

2. 0.0659

2. 0.0659

3. 0.2601

3. 0.2601

4. Not enough information to calculate

4. 信息不足以计算

86\.

86.

What is the relative frequency of direct hits that were AT MOST a category 3 storm?

属于至多 3 级风暴的直袭相对频数是多少?

1. 0.3480

1. 0.3480

2. 0.9231

2. 0.9231

3. 0.2601

3. 0.2601

4. 0.3370

4. 0.3370

1.4 Experimental Design and Ethics 1.4 实验设计与伦理

87\.

87.

How does sleep deprivation affect your ability to drive? A recent study measured the effects on 19 professional drivers. Each driver participated in two experimental sessions: one after normal sleep and one after 27 hours of total sleep deprivation. The treatments were assigned in random order. In each session, performance was measured on a variety of tasks including a driving simulation.

睡眠剥夺如何影响你的驾驶能力?最近一项研究测量了其对 19 名职业司机的影响。每名司机参加了两次实验环节:一次在正常的睡眠之后,一次在总共 27 小时睡眠剥夺之后。处理以随机顺序分配。在每次环节中,通过包括驾驶模拟在内的多种任务测量其表现。

Use key terms from this module to describe the design of this experiment.

运用本模块的关键术语来描述这项实验的设计。

88\.

88.

An advertisement for Acme Investments displays the two graphs in Figure 1.14 to show the value of Acme’s product in comparison with the Other Guy’s product. Describe the potentially misleading visual effect of these comparison graphs. How can this be corrected?

Acme 投资公司的一则广告在图 1.14 中展示了两幅图,以显示 Acme 的产品相对于"另一家"产品的价值。描述这些比较图可能产生的误导性视觉效果。如何对此进行纠正?

89.

89.

The graph in Figure 1.15 shows the number of complaints for six different airlines as reported to the US Department of Transportation in February 2013. Alaska, Pinnacle, and Airtran Airlines have far fewer complaints reported than American, Delta, and United. Can we conclude that American, Delta, and United are the worst airline carriers since they have the most complaints?

图 1.15 中的图显示了 2013 年 2 月向美国运输部报告的六家不同航空公司的投诉数量。Alaska、Pinnacle 和 Airtran 航空公司的报告投诉数远少于 American、Delta 和 United。我们能否因为这些公司投诉最多,就断定 American、Delta 和 United 是最差的航空公司?

Bringing It Together: Homework 综合练习:作业

90. Seven hundred and seventy-one distance learning students at Long Beach City College responded to surveys in the 2010-11 academic year. Highlights of the summary report are listed in Table 1.39.

七百七十一名长滩城市学院的远程学习学生在 2010–11 学年参与了调查。总结报告的要点列于表 1.39。
Table 1.39 LBCC Distance Learning Survey Results
Have computer at home96%
Unable to come to campus for classes65%
Age 41 or over24%
Would like LBCC to offer more DL courses95%
Took DL classes due to a disability17%
Live at least 16 miles from campus13%
Took DL courses to fulfill transfer requirements71%
表 1.39 LBCC 远程学习调查结果
家里有电脑96%
无法到校上课65%
41 岁或以上24%
希望 LBCC 提供更多远程学习课程95%
因残疾而选修远程学习课程17%
住在距校园至少 16 英里以外13%
选修远程学习课程以满足转学要求71%

1. What percent of the students surveyed do not have a computer at home?

1. 被调查的学生中,家里没有电脑的占百分之几?

2. About how many students in the survey live at least 16 miles from campus?

2. 调查中大约有多少学生住在距校园至少 16 英里以外?

3. If the same survey were done at Great Basin College in Elko, Nevada, do you think the percentages would be the same? Why?

3. 如果在内华达州埃尔科的 Great Basin College 进行同样的调查,你认为各项百分比会相同吗?为什么?

91. Several online textbook retailers advertise that they have lower prices than on-campus bookstores. However, an important factor is whether the Internet retailers actually have the textbooks that students need in stock. Students need to be able to get textbooks promptly at the beginning of the college term. If the book is not available, then a student would not be able to get the textbook at all, or might get a delayed delivery if the book is back ordered.

几家在线教材零售商宣称其价格低于校内书店。然而,一个重要因素是这些网络零售商是否确实备有学生所需的教材现货。学生需要在学期初及时取得教材。若教材缺货,学生可能根本无法拿到该教材,或在补货期间遭遇延迟发货。

A college newspaper reporter is investigating textbook availability at online retailers. He decides to investigate one textbook for each of the following seven subjects: calculus, biology, chemistry, physics, statistics, geology, and general engineering. He consults textbook industry sales data and selects the most popular nationally used textbook in each of these subjects. He visits websites for a random sample of major online textbook sellers and looks up each of these seven textbooks to see if they are available in stock for quick delivery through these retailers. Based on his investigation, he writes an article in which he draws conclusions about the overall availability of all college textbooks through online textbook retailers.

一名大学报纸记者正在调查在线零售商的教材可得性。他决定针对以下七个学科各调查一本教材:微积分、生物学、化学、物理学、统计学、地质学与通用工程。他查阅教材行业的销售数据,并在每个学科中选取全国使用最广泛的教材。他访问了若干主要在线教材销售商的网站(随机样本),逐一查询这七本教材,看其是否有现货可通过这些零售商快速发货。基于调查,他撰写了一篇文章,对通过在线教材零售商获取所有大学教材的总体可得性得出结论。

Write an analysis of his study that addresses the following issues: Is his sample representative of the population of all college textbooks? Explain why or why not. Describe some possible sources of bias in this study, and how it might affect the results of the study. Give some suggestions about what could be done to improve the study.

请就该研究撰写一份分析,回答以下问题:他的样本是否能代表所有大学教材这一总体?说明能或不能的理由。描述该研究可能存在的若干偏倚来源,以及它们如何影响研究结果。就如何改进该研究提出建议。