Homeschool Guide: These lesson plans are a guide for parents. Content may contain errors — always cross-reference with official exam board specifications.

data cleaning & reliability

FoundationHigherAll Boards

4 detailed 50-minute lessons with teaching scripts, worked examples, parent guides, and assessment criteria.

Fastmail

Lesson Overview

Total Lessons: 4
Tier: Foundation and Higher
Duration: 50 minutes per lesson (200 minutes total)
Exam Boards: AQA, Edexcel, OCR, Eduqas, CCEA

Learning Objectives

Prerequisites

Materials & Equipment

Lesson 1: Introduction: data cleaning & reliability

Duration: 50 minutes

Starter Activity (5 minutes)

Quick Recall

Write down everything you already know about data cleaning & reliability. Then check against the key terms: key terms from data cleaning & reliability. Use a mini-whiteboard or paper.

Main Content (35 minutes)

Parent/Teacher Guide:
Before lesson: Read the script below. Pre-teach key vocab: key terms from data cleaning & reliability.
If stuck: Re-read the revision notes (link above), then break the content into smaller steps.
Extension: See the Stretch & Challenge ideas in Lesson 4.
Teaching Script (35 mins):
Mins 0-5 - Hook: "Today: data cleaning & reliability. By the end you will be able to answer exam questions on it unaided. It connects to the rest of Statistics because the ideas here recur across the spec."
Mins 5-20 - Direct Instruction: Work through the core ideas below one at a time; after each, ask your student to explain it back in their own words.
Mins 20-30 - Guided Practice: Model the worked example together, then let your student attempt the first practice question with guidance.
Mins 30-35 - Independent Practice: 2-3 practice questions from Lesson 3 below, with immediate feedback.
First Look

Start with the revision notes summary, then attempt: explain the key ideas of data cleaning & reliability

Plenary (5 minutes)

Check Out

Your student states one thing they learned and one question they still have about data cleaning & reliability.

Lesson 2: Core Concepts: data cleaning & reliability

Duration: 50 minutes

Starter Activity (5 minutes)

Review Previous Lesson

Quick recap: write 3 key points from Lesson 1 on data cleaning & reliability. Check them against the notes below.

Main Content (35 minutes)

Key Fact: Data cleaning means checking and correcting a data set before analysis — removing errors, dealing with missing values and identifying anomalies.
Key Fact: Missing data can occur when respondents skip questions or when measurements fail.
Key Fact: Options for missing data: exclude the record, use the mean/median for that variable, or investigate why it is missing.
Key Fact: An outlier is a value that is much higher or lower than the rest of the data.
Key Fact: Outliers should be investigated — they may be genuine extreme values or data entry errors.
Key Fact: If an outlier is a clear error (e.g. height recorded as 25 m), it should be corrected or removed.

Practice (10 minutes)

Q: explain the key ideas of data cleaning & reliability

Answer:

Plenary (5 minutes)

Explain Back

Your student teaches the key points back to you without looking. Fill any gaps immediately.

Lesson 3: Application: data cleaning & reliability

Duration: 50 minutes

Starter Activity (5 minutes)

Quick Recall

Recall the key terms: key terms from data cleaning & reliability. Define each in one sentence.

Main Content (35 minutes)

Parent/Teacher Guide: Let your student attempt each question alone first, then compare with the model answer. Award method marks for correct working even if the final answer is wrong.

Work through the practice questions on the revision notes page for this topic.

Plenary (5 minutes)

Error Review

Review any questions answered incorrectly. Identify whether the error was knowledge, method, or reading the question.

Lesson 4: Exam Practice: data cleaning & reliability

Duration: 50 minutes

Starter Activity (5 minutes)

Command Words

Review what these command words require: state (one point), describe (say what happens), explain (say why), compare (both sides), evaluate (judgement).

Main Content (35 minutes)

Extended Answer

Extended question: Full-Mark Response A data set of students' heights in cm includes the values: 152, 148, 165, 15, 170, 155, 163, 159, 172, 145. Explain how you would clean this data set. <div class="

Step 1 — Identify anomalies: The value 15 cm is clearly an outlier. A height of 15 cm for a student is not realistic, so this is almost certainly a data entry error (likely meant to be 150 or 155). Step 2 — Investigate the outlier: Check the original data collection sheet. If the correct value can be found, correct it. If not, the value should be removed because it is clearly an error and would distort summary statistics (e.g. making the mean far too low). Step 3 — Check for missing data: Confirm all 10 records have been entered. If any records are incomplete, decide whether to exclude them or use an appropriate method to handle the missing values. Step 4 — Verify remaining values: The other heights (145–172 cm) are plausible for students, so they should be kept. After cleaning, the data set should be analysed and the removal of the error value should be documented.

Exam Tips: Always investigate an outlier before removing it — never just delete data without a reason. | When asked about data quality, comment on both reliability and validity. | Bias can occur at any stage: sampling, question wording, data collection or analysis. | Control groups are specific to experiments — do not mention them for surveys or observational studies. | Explain the difference between a genuine outlier and an error using examples.
Common Errors: ✗ Automatically removing all outliers from a data set ✓ Outliers should be investigated first — only remove them if they are clearly errors; genuine extreme values should usually be kept. ✗ Confusing reliability and validity ✓ Reliability is about consistency (repeat results); validity is about whether you are measuring the right thing. A method can be reliable without being valid. ✗ Thinking missing data should always be filled in with the mean ✓ Filling in with the mean can distort results — consider whether the missing data is random or systematic and whether excluding the record is better. ✗ Assuming a large sample guarantees unbiased results ✓ A large sample can still be biased if the s
Stretch & Challenge (Grade 8-9):
  • Synoptic links: explain how data cleaning &amp; reliability connects to another Statistics topic you have studied
  • Real-world: research one real-world use or example of data cleaning &amp; reliability
  • Critical: "What are the limitations of the models used in data cleaning &amp; reliability?"

Plenary (5 minutes)

Assessment Criteria
  • Got it: Confident explanation + correct worked examples
  • Getting there: Main points OK, needs support with detail
  • Not yet: Confused on key concepts - re-run Lesson 2

Homework & Consolidation

Recommended Resources

🎓 Smart Lesson (Guided)