Sign in to view assessments and invite other educators
Sign in using your existing Kendall Hunt account. If you don’t have one, create an educator account.
Help us improve by sharing suggestions or reporting issues.
To Gather
Statistical technology
The mathematical purpose of this activity is for students to investigate the effect of outliers on measures of center and variability, and to make decisions about whether or not to include outliers in a data set.
Arrange students in groups of 2. Provide access to devices that can run GeoGebra or other statistical technology.
Display the data showing Per Capita Health Spending by Country in 2016 for all to see. Orient students to the data, and explain that the distribution of this data set is represented in the histogram used in the Warm-up. Ensure that students understand what per capita means. “Per capita health spending” means the average health spending per person. For example, the United States spends approximately $9,892 on healthcare for each person in the population.
Remind students that we will classify a value in a data set as an outlier if it is greater than Q3 + 1.5
Here is the data set used to create the histogram and box plot from the Warm-up.
Students may incorrectly compute the expression for outliers. Remind them to use the correct order of operations, using the Math Talk Warm-up in a previous lesson.
The goal is to make sure that students understand that outliers can significantly affect measures of center and variability. Discuss the effect of the outlier on the median, mean, and standard deviation and the student responses to “Do you think that 9.8923 should be eliminated from the data set? Why or why not?”
If time permits, discuss questions such as:
Arrange students in groups of 2. Give students quiet think time to answer the first question. Ask partners to compare answers.
In a science class, 11 groups of students are synthesizing biodiesel. At the end of the experiment, each group recorded the mass in grams of the biodiesel they synthesized. The masses of biodiesel are
The purpose of this discussion is to highlight different reasons that outliers appear in data. For example, they could be data-entry or data collection errors, or they could be representative of the sample. The goal is to make sure that students understand that the inclusion of outliers in a data set needs to be evaluated in the context of the data. For the number cube rolls, it is clear that the data should not be used since it is impossible to achieve in the right circumstances. For the other two scenarios, students should understand that a deeper investigation should be done to determine whether the outlier should be included and be able to state circumstances for including or excluding the outlier in each context.
Here are some questions for discussion: