S6 Scatter Graphs

Statistics AQA
GCSE Revision Aid: This resource is designed to support your revision and may contain errors. If you find a discrepancy with your class teaching, your teacher is correct β€” please let us know at gcserevise@scott.scottrix.co.uk.

S6: Scatter Graphs

Foundation Higher AQAEdexcelOCREduqasCCEA

Use and interpret scatter graphs; correlation; lines of best fit; interpolation and extrapolation

Fastmail

πŸ“‹ Key Concepts

Definition: A scatter graph plots two variables against each other to show the relationship between them. Each point represents one data item.

Key Terms

TermDefinition
CorrelationThe relationship between two variables
Line of Best FitA straight line through the data showing the trend
InterpolationPredicting within the range of data
ExtrapolationPredicting beyond the range of data
OutlierA point that doesn't fit the general pattern

πŸ“ Types of Correlation

Positive Correlation: As one variable increases, the other also increases.

Negative Correlation: As one variable increases, the other decreases.

No Correlation: No clear relationship between the variables.
Strength of Correlation:
  • Strong: Points closely follow a line
  • Weak: Points are more scattered but still show a trend
  • None: Points are randomly scattered
Example 1

Describe the correlation you would expect between:

a) Height and weight
b) Price and demand
c) Shoe size and IQ

Solution:

a) Positive correlation - taller people tend to weigh more

b) Negative correlation - higher prices usually mean lower demand

c) No correlation - these are unrelated

πŸ“ Drawing a Line of Best Fit

Rules for Line of Best Fit:
  • Use a ruler to draw a straight line
  • The line should follow the general trend
  • Approximately equal number of points above and below the line
  • The line does NOT have to go through the origin
  • Ignore any outliers when drawing the line
Example 2

A scatter graph shows temperature (x-axis) and ice cream sales (y-axis). The data shows a positive correlation. How would you draw a line of best fit?

Solution:

Draw a straight line that:

  • Has a positive gradient (slopes up from left to right)
  • Has roughly equal points above and below
  • Passes through the "middle" of the data cloud

πŸ“ Using Lines of Best Fit

Interpolation: Using the line to predict a value WITHIN the range of data.

Extrapolation: Using the line to predict a value OUTSIDE the range of data (less reliable).
Example 3

A scatter graph shows revision hours (0-20) and test scores (30-95). The line of best fit passes through (5, 45) and (15, 80).

a) Estimate the score for 10 hours revision.
b) Estimate the score for 25 hours revision.

Solution:

a) 10 hours is within the range (interpolation).

Read up from 10 on x-axis to the line, then across to y-axis.

Estimated score β‰ˆ 62.5

b) 25 hours is outside the range (extrapolation).

Extend the line and read off: β‰ˆ 97.5

Less reliable - we don't know if the pattern continues.

Warning about Extrapolation: Predictions outside the data range may be unreliable because the relationship might not continue in the same way.

πŸ“ Identifying Outliers

Outlier: A data point that is significantly different from the general pattern. It may be an error or a genuine unusual case.
Example 4

A scatter graph shows age and salary. Most points follow a trend, but one point shows a 22-year-old earning Β£200,000. Is this an outlier?

Solution:

Yes, this point is an outlier. It is far from the general pattern where salary increases with age. This could be:

  • A data entry error (perhaps Β£20,000?)
  • A genuine case (e.g. a successful entrepreneur or athlete)

Outliers should be investigated - don't just delete them without checking.

πŸ“ Causation vs Correlation

Important: Correlation does NOT prove causation!

Just because two variables are correlated, it doesn't mean one causes the other.
Example 5

Studies show positive correlation between ice cream sales and drowning deaths. Does ice cream cause drowning?

Solution:

No! Both are caused by a third variable: hot weather.

Hot weather β†’ More people buy ice cream

Hot weather β†’ More people go swimming β†’ More drowning

This is called a "spurious correlation" or "confounding variable".

Possible explanations for correlation:
  • Variable A causes B
  • Variable B causes A
  • A third variable causes both A and B
  • It could be a coincidence

πŸ“ Equation of Line of Best Fit

Form: y = mx + c
where m = gradient (change in y Γ· change in x)
and c = y-intercept (where line crosses y-axis)
Example 6

A line of best fit passes through (2, 10) and (8, 40). Find the equation.

Solution:

Gradient = (40 - 10) Γ· (8 - 2) = 30 Γ· 6 = 5

y = 5x + c

Substitute (2, 10): 10 = 5(2) + c, so c = 0

Equation: y = 5x

Example 7

Using the line y = 5x, predict y when x = 6.

Solution:

y = 5 Γ— 6 = 30

❓ Practice Questions

Q1: What type of correlation would you expect between hours of sleep and tiredness?

Q2: A line of best fit has equation y = 3x + 5. Predict y when x = 7.

Q3: What is the difference between interpolation and extrapolation?

Q4: A scatter graph shows test scores between 40% and 90%. Is predicting a score of 95% interpolation or extrapolation?

Q5: State two rules for drawing a line of best fit.

βœ… Answers

  1. Negative correlation - more sleep means less tiredness.
  2. y = 3(7) + 5 = 21 + 5 = 26
  3. Interpolation is predicting within the data range; extrapolation is predicting outside the range.
  4. Extrapolation - 95% is outside the range of 40-90%.
  5. Any two: use a ruler; follow the trend; equal points above and below; don't force through origin; ignore outliers.

🎯 Exam Tips

🧠 Problem-Solving Strategies

Problem-Solving

For scatter graph problems: (1) Describe correlation using "positive", "negative", or "none" β€” never "good" or "bad", (2) Draw lines of best fit with roughly equal points above and below, (3) For predictions, use the line of best fit β€” not individual points, (4) Interpolation (within data range) is more reliable than extrapolation (outside), (5) Correlation does NOT prove causation β€” always consider confounding variables, (6) Find the equation of the line of best fit using y = mx + c.
Multi-Step Problem

A scatter graph shows revision hours (x) and test scores (y). The line of best fit passes through (5, 40) and (20, 85). (a) Find the equation of the line of best fit. (b) Predict the score for 12 hours revision. (c) Explain why predicting the score for 30 hours may be unreliable.

Solution: (a) Gradient = (85βˆ’40)/(20βˆ’5) = 45/15 = 3. y βˆ’ 40 = 3(x βˆ’ 5), so y = 3x + 25. (b) y = 3(12) + 25 = 61. (c) 30 hours is outside the data range (0–20 hours), so this is extrapolation. The linear relationship may not continue β€” students might reach a maximum score or tire out, so the prediction may be too high.

⚠️ Common Errors

Watch Out!

1. Wrong: Saying "strong positive correlation" means one variable causes the other Correct: Correlation does NOT imply causation β€” a third variable could cause both, or it could be coincidence

2. Wrong: Drawing a line of best fit that passes through (0,0) because "it should start at the origin" Correct: The line of best fit does NOT have to go through the origin β€” it should follow the trend of the data with roughly equal points above and below

3. Wrong: Describing correlation as "good" or "bad" instead of positive/negative/none Correct: Use "positive correlation" (both increase), "negative correlation" (one increases, other decreases), or "no correlation" β€” avoid value judgements

✍️ 6-Mark Exam Question

Extended Answer

6 marks: Data is collected on ice cream sales and temperature for 12 days:

Temperature (Β°C): 15, 18, 20, 22, 24, 25, 26, 28, 30, 32, 33, 35

Sales (Β£hundreds): 5, 8, 10, 14, 16, 18, 19, 22, 25, 28, 30, 35

(a) Describe the correlation. (b) The line of best fit has equation y = 1.5x βˆ’ 18. Use this to predict sales at 21Β°C and 40Β°C. (c) Which prediction is more reliable? Explain why.

(a) Strong positive correlation β€” as temperature increases, ice cream sales also increase. The points follow a clear upward trend.

(b) At 21Β°C: y = 1.5(21) βˆ’ 18 = 31.5 βˆ’ 18 = 13.5, so Β£1,350. At 40Β°C: y = 1.5(40) βˆ’ 18 = 60 βˆ’ 18 = 42, so Β£4,200.

(c) The prediction at 21Β°C is more reliable because it is interpolation β€” 21Β°C is within the data range (15–35Β°C). The prediction at 40Β°C is extrapolation β€” it's outside the range, and the relationship might not continue linearly. For example, at very high temperatures, people might stay indoors and sales could drop.

Mark scheme: M1 for identifying positive correlation, A1 for "strong positive", M1 for correct substitution, A1 for both predictions, M1 for identifying interpolation vs extrapolation, A1 for explaining why 21Β°C is more reliable

πŸ“Š AO3: Reason & Interpret

Reasoning and Interpretation

A study finds a strong positive correlation between the number of books in a home and children's test scores.

(a) Describe the correlation in context.

(b) A politician says "If we give every family more books, test scores will improve." Evaluate this claim.

(c) Suggest a confounding variable that could explain the correlation.

Answers: (a) Strong positive correlation β€” homes with more books tend to have children with higher test scores. (b) The claim is not well supported β€” correlation does not prove causation. Having more books doesn't necessarily cause higher scores; the relationship might be due to other factors. (c) Income/wealth β€” wealthier families can afford more books AND better education/tutoring. Parental education level is another confounding variable β€” more educated parents may have more books and also provide more academic support.

πŸ“ Exam Questions by Topic

🎬 Video Resources

Share this page

Ready to ace your GCSE Mathematics exams?

Get the best revision books and guides to boost your grades.

🧠 Flashcards (Spaced Repetition)

πŸ“ Exam Questions by Topic

🎯 Target Tests (Auto-Graded)

πŸ“ Exam Questions by Topic

πŸ“„ Past Papers for Statistics (AQA)

For the most accurate and up-to-date past papers, always check the official exam board websites.