Find the shaded area between the function, the -axis, and the boundaries and . Explain or show your reasoning.
What proportion of the area between the function, the -axis, and the boundaries and is shaded? Explain or show your reasoning.
A bell curve, made up of smaller line segments. The horizontal axis is labeled from 0 to 8 by ones. The vertical axis is labeled from 0 to 8 by ones. A line segment from 0 comma 0 to 1 comma 1, a line segment from 1 comma 1 to 2 comma 3, a line segment from 2 comma 3 to 3 comma 4, a line segment from 3 comma 4 to 4 comma 4, a line segment from 4 comma 4 to 5 comma 3, a line segment from 5 comma 3 to 6 comma 1, and a line segment from 6 comma 1 to 7 comma 0. The area under the curve between 1 and 2 on the horizontal axis is shaded.
6.2
Activity
Story Submissions
A publisher takes submissions for short stories to include in a book. 200 stories are submitted, and the publisher needs to be aware of how long each story is. The way the publisher will put together the collection of stories, a page typically contains 200 words. The mean number of words for each story is 2,600, and the standard deviation is 400 words.
If a histogram is created using intervals of 200 words, what would be the area of the bar representing the number of stories that contain between 2,000 and 2,200 words? Explain or show your reasoning.
What proportion of the total area is represented by the bar for stories that contain between 2,000 and 2,200 words? Explain or show your reasoning.
What proportion of stories in this group contains between 2,000 and 2,200 words? Explain or show your reasoning.
How does the proportion of the area you calculated relate to the proportion of stories in the group that contain between 2,000 and 2,200 words?
What proportion of stories in this group is within 1 standard deviation of the mean number of words?
What proportion of stories in this group is within 2 standard deviations of the mean number of words?
What proportion of stories in this group is within 1 standard deviation of 2,400 words?
6.3
Activity
Website Load Times
A company collects data from 10,000 websites about how long it takes to load the site. The number of seconds it takes to fully load the website is summarized in the relative frequency table.
seconds
to load
relative frequency
1.4–1.6
0.0003
1.6–1.8
0.0012
1.8–2.0
0.0053
2.0–2.2
0.0181
2.2–2.4
0.0442
2.4–2.6
0.0910
2.6–2.8
0.1555
2.8–3.0
0.1861
3.0–3.2
0.1938
3.2–3.4
0.1447
3.4–3.6
0.0923
3.6–3.8
0.0447
3.8–4.0
0.0166
4.0–4.2
0.0048
4.2–4.4
0.0012
4.4–4.6
0.0002
The relative frequency histogram summarizes the same data.
Histogram. The vertical axis is labeled from 0 to .18 by 0.02’s. The horizontal axis is labeled seconds to load, and has bin widths of 0.2, starting at 1.4. Beginning at 1.4 up to, but not including 1.6, height of bar at each interval is 0.0003, ,0.0012, 0.0053, 0.0181, 0.0442, .0910, .1555, .1861, 0.1938, .1447, .0923, .0447, .0166, .0048, .0012, .0002, 0.
The mean time to load a website is 3 seconds, and the standard deviation is 0.4 second.
Would a normal distribution be a good model for this distribution? Explain your reasoning.
What proportion of websites loaded within 1 standard deviation of the mean?
What proportion of websites loaded within 2 standard deviations of the mean?
What proportion of websites loaded within 1 standard deviation of 2.8 seconds?
Compare the proportion of websites within 1 standard deviation of the mean to the proportion of stories in the submissions that are within 1 standard deviation of the mean number of words from the previous task. Do the same for the proportion within 2 standard deviations.
Student Lesson Summary
There is an important connection between areas in histograms and the data represented by the histogram. In particular, the proportion of the total area in the histogram that is represented by a single bar in the histogram is equivalent to the proportion of all the data that is included in that interval. This is made more interesting by the fact that, for normally distributed data, the proportion of values in an interval whose endpoints are described by the mean and standard deviation is always the same.
For example, a woodshop produces boards of various lengths. During a certain week, 5,000 boards are produced and measured. The mean length is 6 feet, and the standard deviation length is 1 foot. The table and histogram show a summary of the board lengths.
board length
3.5–4
4–4.5
4.5–5
5–5.5
5.5–6
6–6.5
6.5–7
7–7.5
7.5–8
8–8.5
frequency
113
220
460
747
960
955
753
460
220
112
Histogram. The vertical axis is labeled from 0 to 1000 by one hundred. The horizontal axis is labeled board length, feet, and has bin widths of 0.5, starting at 3. Beginning at 3 up to, but not including 3.5, height of bar at each interval is 0, 113, 220, 460, 747, 960, 955, 753, 460, 220, 112, 0.
The total area of all the rectangles in the histogram is 2,500 since we could stack all the bars on top of one another and have a rectangle that is 5,000 tall and 0.5 wide. If we look at just the rectangles representing boards between 5.5 and 6 feet wide, the area is 480, which is 19.2% of the total area, since . Similarly, we can see from the data that 19.2% of the data is in this same interval since . It is not a coincidence that these values are the same! The proportion of the total area that is in one of the rectangles is always equivalent to the proportion of all the data values that are in the same interval.
When the data is normally distributed, the proportions of certain regions are always the same. For example, there is always about 68% of the data within 1 standard deviation of the mean. Since the boards produced by the woodshop are approximately normal, we can test this information.
The boards within one standard deviation of the mean are between 5 and 7 feet long. Using the table, we can see that 3,415 boards are in this range () and those represent 68.3% () of the boards produced in the woodshop.
Let’s say that, another week, the woodshop produces 5,000 boards again, but this time, the mean is 6.5 feet and the standard deviation is 0.75 foot. As long as the board lengths continue to be approximately normal, we can expect about 68% of the boards to be within 1 standard deviation of the mean. For that week, it means that about 68% of the boards are between 5.75 and 7.25 feet long.
In fact, as long as the interval can be described using only the mean and standard deviation, and the data is normally distributed, the proportion of data values in the interval can be found. In general, about 68% of the data is within 1 standard deviation of the mean, about 95% of the data is within 2 standard deviations of the mean, and more than 99% of the data is within 3 standard deviations of the mean.
Glossary
None
Have feedback on the curriculum?
Help us improve by sharing suggestions or reporting issues.