Populations, Samples and Sampling Methods

Populations, Samples and Sampling Methods

Population vs sample

Introduction

You rarely have access to every single data point you care about. Measuring everyone is expensive, slow, or simply impossible. That is why statistics works with samples - subsets of the population - and uses them to make statements about the whole.

Population vs Sample

  • Population: every individual or item of interest. Example: all 2 million customers of a webshop.
  • Sample: the subset you actually collect data from. Example: 400 randomly chosen customers.
Population vs sample

A parameter (like population mean μ) is what you want to know. A statistic (like sample mean x̄) is what you can compute. The gap between them is sampling error - not a mistake, just the unavoidable price of not measuring everyone.

Why Not Just Measure Everything?

SituationWhy sampling wins
2 million customersSurveying all is too expensive
Quality testing light bulbsTesting destroys the product
Measuring blood pressure of patientsInvasive, impractical for all
Tracking website usersImpossible to catch every session

Sampling Methods

1. Simple Random Sampling

Every member of the population has an equal chance of being selected.

import numpy as np

customers = np.arange(1, 2001) # 2000 customers sample = np.random.choice(customers, size=100, replace=False) print(sample[:10])

Output:

[ 421 1856 1410  902  332 1779  611  743 1296 1154]

2. Stratified Sampling

Split the population into groups (strata), then draw randomly from each group. This guarantees representation.

# Customers split by region; sample 50 from each region
regions = {"PL": np.arange(1, 1001), "DE": np.arange(1001, 1601), "FR": np.arange(1601, 2001)}
sample = np.concatenate([np.random.choice(v, 50, replace=False) for v in regions.values()])
print(f"stratified sample size: {len(sample)}")

3. Systematic Sampling

Pick every k-th element from an ordered list.

population = np.arange(1, 2001)
k = 20
sample = population[::k]   # every 20th customer
print(len(sample))

4. Cluster Sampling

Divide the population into clusters, randomly choose some clusters, then measure everyone inside them. Cheap but less precise.

Bias: The Silent Killer

A sample is only useful if it represents the population. Bias makes a sample systematically different from the population:

Bias typeExample
Selection biasSurveying only active app users
Response biasOnly happy customers reply to feedback
Survivorship biasAnalyzing only companies that survived
Convenience biasAsking friends instead of random users

Example: The Classic Mistake

In 1948 the Chicago Tribune predicted Dewey would beat Truman in the US election, based on a huge sample of over 50,000 phone and mail surveys. The sample was enormous but biased: only people with phones (wealthier) were reached. Truman won. Big sample size does not fix bias.

How Large Should a Sample Be?

Larger samples shrink sampling error, but the gain shrinks too. The standard error of the mean is:

SE = sigma / sqrt(n)

Doubling the sample size reduces the standard error by only about 30% (factor 1/sqrt(2)). Ten times more data gives you sqrt(10) ≈ 3.2 times more precision.

import numpy as np

for n in [50, 200, 800, 3200]: ses = [] for _ in range(2000): s = np.random.normal(100, 15, n) ses.append(s.std(ddof=1) / np.sqrt(n)) print(f"n = {n:5d} avg SE = {np.mean(ses):.2f}")

Output:

n =    50  avg SE = 2.12
n =   200  avg SE = 1.06
n =   800  avg SE = 0.53
n =  3200  avg SE = 0.27

Notice how quadrupling n only halves the standard error.

Summary

  • A sample is a subset used to estimate population parameters.
  • Statistics (sample) estimate parameters (population).
  • Random, stratified, systematic and cluster sampling each have trade-offs.
  • Bias cannot be cured by a bigger sample.
  • Standard error decreases with sqrt(n).

Next Lesson

In the next lesson you will start descriptive statistics: measures of central tendency - the mean, median and mode, and when to use each.

Quiz - Quiz - Populations and Samples

1. A parameter is a number that describes:

2. Which sampling method splits the population into groups and draws randomly from each?

3. In the 1948 election poll mistake, the huge sample was biased because:

4. Does a very large biased sample give accurate results?

5. If you quadruple the sample size, the standard error:

Types of Data and Measurement Scales