Sampling, estimation and uncertainty

A sample proportion is a moving target.

Distinguish a population probability from a sample estimate, derive sampling variability, and identify bias that more data cannot remove.

Start with: Probability and averages. A Bernoulli trial has two outcomes with fixed success probability p.

01 · Commit to a prediction

What do you expect?

Responses stay in this page only. Reloading or closing may discard them. Nothing is transmitted, saved or synchronized.

02 · Change an assumption

Predict. Change. Explain.

Before changing a control, say what should move and why. Start with the experiments below. Reset restores the starting model; it preserves your written responses.

Experiment 1

Keep p = 50%. Change n from 4 to 16. Predict the standard error, not the height of the middle bar.

Check the prediction

Standard error halves from 25 to 12.5 percentage points. The graph now has more possible proportions; one bar need not gain probability.

Experiment 2

Reset. Set p to 0%, then 100%. What randomness remains?

Check the prediction

None under this fixed model. Every trial has the same outcome. Real uncertainty about an unknown p is a different question.

03 · Connect the mechanism

From the picture to the quantities.

The population parameter p stays fixed while a sample estimate p̂ = X/n varies across hypothetical repetitions. The graph enumerates every possible success count with its exact probability. It draws no random sample. Standard error describes the spread of the estimator, not the spread of individual binary outcomes and not the error of a particular sample.

P(X=k) = C(n,k) pᵏ (1−p)ⁿ⁻ᵏ
E[p̂] = p
SE(p̂) = √[p(1−p)/n]

n is the number of independent trials, 1 to 100 here. X is the integer success count and k one possible count, 0 to n. p and p̂ are fractions; displayed proportions use percent. C(n,k) counts arrangements. E means expectation across repeated samples. SE has the same units as p̂. Endpoint distributions are defined directly, without relying on ambiguous 0⁰ notation.

Worked example

For n = 4 and p = 1/2, counts 0,1,2,3,4 have probabilities 1,4,6,4,1 divided by 16. Their mean proportion is 1/2. The squared deviations of proportions are 1/4,1/16,0,1/16,1/4. Weighting and adding gives variance 1/16, so the standard error is 1/4 = 25 percentage points.

Assumptions and limits

This is a known-p thought experiment, not a confidence interval fitted to observations. Independence and a common p are essential. Finite-population sampling without replacement, clustered observations and changing probabilities need different models. More trials shrink random error under the model; they cannot fix selection bias or incorrect measurement.

The annotated sources distinguish established results from this lesson’s original examples.

04 · Follow the structure

Where else does this apply?

Quality sampling

Independent items from a stable process produce variable defect estimates.

Boundary: Batches can induce dependence and production conditions can drift.

Randomized surveys

A random sample can estimate a defined population fraction.

Boundary: Voluntary-response polls may be biased even when enormous.

05 · Retrieve without hints

Close the explanation. Try a new case.

Write an answer before opening its feedback. Later, return directly here without rereading above. Recognition, explanation and transfer are separate outcomes. No page action or answer reveal measures mastery.

Recognition and explanation

Reveal reasoning and rubric

No. It is a standard deviation across repeated samples. Individual deviations can be larger, and bias is outside this model.

Self-check: Distinguish spread, coverage and bias.

Calculation

Reveal reasoning and rubric

n = 25 gives √(0.25/25) = 0.1. n = 100 gives 0.05. Halving standard error requires four times the sample size.

Self-check: Use proportions in the equation and convert to percentage points at the end.

Novel transfer

Reveal reasoning and rubric

No. Repeated images of the same binary item share its outcome. With perfect duplicates the effective independent count remains 100. The model must represent clustering rather than count files as trials.

Self-check: Identify the independent unit and why duplicated measurements add no new outcome information.

On a later day, try again and record actual evidence in the curriculum. A later unaided explanation and a fresh transfer problem give stronger evidence than immediate familiarity. No reminder is scheduled.

Sources & scope

What supports the lesson?

Original teaching examples. Reference links need a connection; the lesson itself does not. Built 2026-10-11. Learner understanding is not assessed.