ST8: Data Cleaning & Reliability
Understand how to clean data by dealing with missing values and outliers, and assess data quality through reliability, validity, bias and control groups.
Understand how to clean data by dealing with missing values and outliers, and assess data quality through reliability, validity, bias and control groups.
Understand how to clean data by dealing with missing values and outliers, and assess data quality through reliability, validity, bias and control groups.
For Data Cleaning & Reliability, you must know:
Q: What is data cleaning?
Q: How should you deal with an outlier?
Q: What is the difference between reliability and validity?
Q: Why is a control group important in an experiment?
Q: Give an example of bias in data collection.
✗ Automatically removing all outliers from a data set ✓ Outliers should be investigated first — only remove them if they are clearly errors; genuine extreme values should usually be kept.
✗ Confusing reliability and validity ✓ Reliability is about consistency (repeat results); validity is about whether you are measuring the right thing. A method can be reliable without being valid.
✗ Thinking missing data should always be filled in with the mean ✓ Filling in with the mean can distort results — consider whether the missing data is random or systematic and whether excluding the record is better.
✗ Assuming a large sample guarantees unbiased results ✓ A large sample can still be biased if the sampling method is flawed (e.g. an online survey excludes people without internet access).
A data set of students' heights in cm includes the values: 152, 148, 165, 15, 170, 155, 163, 159, 172, 145. Explain how you would clean this data set.
Step 1 — Identify anomalies: The value 15 cm is clearly an outlier. A height of 15 cm for a student is not realistic, so this is almost certainly a data entry error (likely meant to be 150 or 155). Step 2 — Investigate the outlier: Check the original data collection sheet. If the correct value can be found, correct it. If not, the value should be removed because it is clearly an error and would distort summary statistics (e.g. making the mean far too low). Step 3 — Check for missing data: Confirm all 10 records have been entered. If any records are incomplete, decide whether to exclude them or use an appropriate method to handle the missing values. Step 4 — Verify remaining values: The other heights (145–172 cm) are plausible for students, so they should be kept. After cleaning, the data set should be analysed and the removal of the error value should be documented.
AO1 (Knowledge & Understanding): Demonstrate knowledge and understanding of data cleaning & reliability, including data collection, presentation and calculation techniques relevant to Edexcel 1ST0 & AQA 8382.
AO2 (Application): Apply knowledge and understanding of data cleaning & reliability to interpret data, reason statistically and draw conclusions in context.
AO3 (Evaluation): Evaluate statistical methods and conclusions, assessing appropriateness, reliability, validity and bias through the statistical enquiry cycle.
Get the best revision books and guides to boost your grades.
For the most accurate and up-to-date past papers, always check the official exam board websites.