ST16: Skewness & Outliers
Learn how to identify positive and negative skew, calculate Pearson's skewness coefficient, detect outliers using 1.5xIQR and 2 standard deviations, and understand the effect of outliers on averages.
Learn how to identify positive and negative skew, calculate Pearson's skewness coefficient, detect outliers using 1.5xIQR and 2 standard deviations, and understand the effect of outliers on averages.
Learn how to identify positive and negative skew, calculate Pearson's skewness coefficient, detect outliers using 1.5xIQR and 2 standard deviations, and understand the effect of outliers on averages.
For Skewness & Outliers, you must know:
Q: For a distribution with mean = 45, median = 40, and SD = 10, calculate Pearson's skewness coefficient and interpret it.
Q: A dataset has Q1 = 20, Q3 = 44. A value of 82 is recorded. Is it an outlier using the 1.5 ร IQR rule?
Q: The data 5, 6, 7, 8, 100 has mean 25.2 and median 7. Why is the median more appropriate?
Q: On a box plot, the right whisker is much longer than the left and the median is near Q1. What does this indicate?
Q: A value of 12 is recorded in a dataset with mean 50 and SD 18. Is it a potential outlier using the 2 SD rule?
โ Saying a distribution is skewed just because it is not perfectly symmetric โ Small deviations from symmetry are normal; only describe a distribution as skewed when the asymmetry is clearly visible or the skewness coefficient is substantial.
โ Confusing the direction of skew โ In positive skew, the tail goes to the RIGHT (towards higher values) and the mean is above the median. In negative skew, the tail goes to the LEFT and the mean is below the median.
โ Forgetting to check whether an outlier is a data error before deciding to remove it โ Always investigate outliers first; if a value is a genuine measurement, it should not be removed without justification. Only remove values confirmed as recording errors.
โ Applying the 2 SD rule to heavily skewed data โ The 2 SD rule relies on the assumption of approximate normality; for skewed data, the 1.5 ร IQR rule is more appropriate.
A dataset of 10 test scores has Q1 = 35, Q2 = 52, Q3 = 65, mean = 54, and SD = 18. The maximum value is 110. Determine whether the maximum is an outlier and calculate Pearson's skewness coefficient.
Outlier test using 1.5 ร IQR: IQR = 65 โ 35 = 30 Upper boundary = 65 + 1.5 ร 30 = 65 + 45 = 110 Since 110 equals (not exceeds) the boundary, it is not classified as an outlier by the 1.5 ร IQR rule. It is at the very edge. Pearson's skewness coefficient: Skewness = 3 ร (mean โ median) รท SD = 3 ร (54 โ 52) รท 18 = 6 รท 18 = 0.333 Since 0.333 is between โ0.5 and 0.5, the distribution is approximately symmetric with a very slight positive skew.
AO1 (Knowledge & Understanding): Demonstrate knowledge and understanding of skewness & outliers, including data collection, presentation and calculation techniques relevant to Edexcel 1ST0 & AQA 8382.
AO2 (Application): Apply knowledge and understanding of skewness & outliers to interpret data, reason statistically and draw conclusions in context.
AO3 (Evaluation): Evaluate statistical methods and conclusions, assessing appropriateness, reliability, validity and bias through the statistical enquiry cycle.
Get the best revision books and guides to boost your grades.
For the most accurate and up-to-date past papers, always check the official exam board websites.