GCSE Revision Aid: This resource is designed to support your revision and may contain errors. If you find a discrepancy with your class teaching, your teacher is correct โ€” please let us know at gcserevise@scott.scottrix.co.uk.

S5: Describing Populations

Foundation Higher AQAEdexcelOCREduqasCCEA

Apply statistics to describe a population

Fastmail

๐Ÿ“‹ Key Concepts

Definition: Describing a population involves using statistical measures to summarise and communicate the key characteristics of a group.

Key Statistics for Description

MeasureWhat It Tells Us
MeanThe typical average value
MedianThe middle value (less affected by outliers)
ModeThe most common value
RangeHow spread out the data is
IQR (Higher)Spread of middle 50%

๐Ÿ“ Writing a Statistical Summary

A good statistical summary should:
  • State the average (mean or median)
  • Describe the spread (range or IQR)
  • Comment on the shape of the distribution
  • Make comparisons where relevant
  • Draw conclusions about the population
Example 1

A survey of 50 households found:

  • Mean number of occupants: 2.8
  • Median: 3
  • Range: 7

Write a summary of this data.

Solution:

"The average household has approximately 3 occupants. The median of 3 suggests that half of households have 3 or more people. The range of 7 indicates considerable variation in household size, from 1 to 8 occupants."

๐Ÿ“ Describing Distribution Shape

Key terms for distribution shape:
  • Symmetrical: Mean โ‰ˆ Median, data balanced around centre
  • Skewed right (positive): Mean > Median, tail extends to higher values
  • Skewed left (negative): Mean < Median, tail extends to lower values
  • Bimodal: Two peaks, two distinct groups in data
Example 2

Test scores have mean = 62, median = 68. What does this tell us about the distribution?

Solution:

Mean (62) < Median (68)

The distribution is skewed left (negative skew).

A few very low scores pull the mean down, but most students scored around 68 or higher.

Example 3

House prices have mean = ยฃ350,000 and median = ยฃ280,000. Describe the distribution.

Solution:

Mean > Median indicates right skew (positive skew).

A few expensive houses pull the mean up above the median.

Most houses are priced around ยฃ280,000, with some much more expensive properties.

๐Ÿ“ Making Inferences

Inference: Using sample statistics to draw conclusions about the population.

Always acknowledge limitations:
  • Sample may not be fully representative
  • Results are estimates, not certainties
  • Consider the sample size and method
Example 4

A random sample of 200 adults found mean height = 171 cm, range = 42 cm. What can we infer about the population?

Solution:

"We can infer that the average adult height is approximately 171 cm. Heights vary considerably with a range of 42 cm. However, this is based on a sample, so the true population mean may differ slightly. A larger sample would give more reliable estimates."

๐Ÿ“ Using Appropriate Statistics

Choosing the right measure:
  • Use mean: When data is roughly symmetrical with no outliers
  • Use median: When data has outliers or is skewed
  • Use mode: For categorical data or to find most common value
Example 5

Salaries at a company: ยฃ18k, ยฃ22k, ยฃ25k, ยฃ24k, ยฃ21k, ยฃ95k, ยฃ23k. Which average should you use?

Solution:

Use the median. The ยฃ95k salary is an outlier (likely the owner/manager).

Mean = (18+22+25+24+21+95+23) รท 7 = 32.6k (misleading)

Median = 23k (better represents typical salary)

๐Ÿ“ Writing Conclusions

Good conclusions:
  • Are based on the statistical evidence
  • Use appropriate language ("suggests", "indicates")
  • Acknowledge limitations
  • Answer the question asked
Example 6

A sample of 100 students shows mean study time = 12 hours/week with range = 18 hours. Write a conclusion about study habits.

Solution:

"Students typically spend about 12 hours per week studying, though there is considerable variation (range of 18 hours). Some students study much more than others. This suggests that study habits vary widely among students, possibly due to different subjects, year groups, or personal circumstances."

๐Ÿ“ Comparing Populations

Example 7

Compare these two schools' test results:

School A: Mean = 68, Range = 35
School B: Mean = 72, Range = 20

Solution:

"School B performed better on average (mean 72 vs 68). School B's results were more consistent with a smaller range (20 vs 35), suggesting more uniform teaching or student ability. School A has more variation in performance, with some students doing much better or worse than others."

โ“ Practice Questions

Q1: Data has mean = 45, median = 38. Describe the skew of the distribution.

Q2: A survey finds mode = 2 for "number of pets". What does this tell us?

Q3: Why might the median be better than the mean for describing house prices?

Q4: Two classes have test results: Class A (Mean=65, Range=40), Class B (Mean=68, Range=15). Compare them.

Q5: What limitations should you mention when describing a population from sample data?

โœ… Answers

  1. Right skew (positive skew) - mean > median suggests a few high values pulling the mean up.
  2. Most households have 2 pets - this is the most common value.
  3. House prices often have expensive outliers that pull the mean up. The median better represents a "typical" price.
  4. Class B performed better on average (68 > 65). Class B has more consistent results (smaller range: 15 vs 40). Class A has more variation in performance.
  5. Sample may not be fully representative; results are estimates; sampling error exists; larger samples give more reliable conclusions.

๐ŸŽฏ Exam Tips

๐Ÿง  Problem-Solving Strategies

Problem-Solving

When describing populations: (1) Always report BOTH an average (mean or median) AND a measure of spread (range or IQR), (2) Compare mean vs median to identify skew โ€” if mean > median, the data is right-skewed, (3) Use cautious language: "suggests", "indicates", "approximately" โ€” sample data gives estimates, not certainties, (4) When comparing two groups, discuss both centre and spread, (5) Acknowledge limitations: sample size, sampling method, possible bias.
Multi-Step Problem

Two towns have house price data: Town A: Mean = ยฃ195,000, Median = ยฃ180,000, IQR = ยฃ60,000. Town B: Mean = ยฃ210,000, Median = ยฃ205,000, IQR = ยฃ40,000. Compare the two towns and describe the distribution in each.

Solution: Town B has higher house prices on average (both mean and median). Town A shows right skew (mean ยฃ195k > median ยฃ180k) โ€” some expensive houses pull the mean up. Town B is more symmetrical (mean โ‰ˆ median). Town A has more variation (IQR ยฃ60k vs ยฃ40k), so prices in Town A are less consistent. Town B has more uniform pricing.

โš ๏ธ Common Errors

Watch Out!

1. Wrong: Only commenting on the average when comparing distributions, ignoring spread Correct: Always compare BOTH average and spread โ€” "Town B is higher on average AND more consistent"

2. Wrong: Making definitive statements like "the population mean is exactly 42" from sample data Correct: Use cautious language โ€” "the sample suggests the population mean is approximately 42"

3. Wrong: Using the mean to describe skewed data without acknowledging it's misleading Correct: For skewed data, the median is more representative โ€” always note the skew and choose the appropriate average

โœ๏ธ 6-Mark Exam Question

Extended Answer

6 marks: A sample of 150 adults in City X reports: mean commute time = 34 minutes, median = 28 minutes, range = 75 minutes. A sample of 120 adults in City Y reports: mean = 31 minutes, median = 30 minutes, range = 40 minutes. (a) Describe the shape of the distribution for each city. (b) Compare the commute times for the two cities. (c) Ravi says "City X has longer commutes than City Y." Evaluate this claim using both averages.

(a) City X: Mean (34) > Median (28) โ€” right-skewed, with some very long commutes pulling the mean up. City Y: Mean (31) โ‰ˆ Median (30) โ€” roughly symmetrical distribution.

(b) City X has slightly higher average commute (mean 34 vs 31, median 28 vs 30). However, the median is actually lower in City X (28 vs 30), suggesting most people in City X have shorter commutes but a few have very long ones. City X has much more variation (range 75 vs 40).

(c) Ravi's claim is partly supported by the mean but contradicted by the median. The mean in City X is higher due to a few extreme values (outliers with very long commutes). The median โ€” a better measure for skewed data โ€” is actually lower in City X. So for a "typical" person, City Y has slightly longer commutes. Ravi's claim is misleading because the mean is distorted by outliers.

Mark scheme: M1 for identifying right skew in X, A1 for symmetrical in Y, M1 for comparing both averages, A1 for noting median reversal, M1 for evaluating the claim, A1 for explaining why mean is misleading

๐Ÿ“Š AO3: Reason & Interpret

Reasoning and Interpretation

A sample of 80 students found: mean number of hours of sleep per night = 7.2, median = 7.5, range = 6. A separate study of 2000 students nationally found mean = 7.8, median = 7.8.

(a) What does the difference between mean and median in the first sample suggest about the distribution?

(b) Can we reliably conclude that these 80 students sleep less than the national average? Give two reasons.

(c) Which study gives a more reliable estimate of the true average? Why?

Answers: (a) Mean (7.2) < Median (7.5) โ€” left-skewed. A few students with very low sleep hours are pulling the mean down. (b) Not reliably โ€” (1) the sample is small (only 80), so sampling error could explain the difference, (2) the sample may not be representative (e.g. all from one school with early start times). (c) The national study (n=2000) is more reliable because a much larger sample size means the sample mean is closer to the true population mean, with less sampling error.

๐Ÿ“ Exam Questions by Topic

๐ŸŽฌ Video Resources

Share this page

Ready to ace your GCSE Mathematics exams?

Get the best revision books and guides to boost your grades.