What is Statistics and Why It Matters

What is Statistics and Why It Matters

Distribution of customer order values

Introduction

Statistics is the science of collecting, organizing, analyzing, and interpreting data to make decisions. For a data analyst, statistics is not an abstract academic subject: it is the toolkit you use every day to answer questions like "is this campaign working?", "which customer group is most valuable?", or "is the difference between two versions real or just random?"

Descriptive vs Inferential Statistics

Statistics splits into two big branches:

  • Descriptive statistics summarize and describe the data you have. Examples: average order value, percentage of returning customers, a histogram of website visit times.
  • Inferential statistics use a sample to make conclusions about a larger population. Examples: estimating the average height of all users from a sample of 500, testing whether a new layout increases conversion.

Example: Order Values

Imagine an e-commerce store. You collect 500 customer orders and build a histogram of order values:

Histogram of order values

The histogram above shows:

  • most orders are between 30 and 60 dollars,
  • a second smaller cluster appears around 70 dollars,
  • the distribution is not perfectly symmetric.
Describing this shape is descriptive statistics. Using those 500 orders to estimate the average order value of all future customers is inferential statistics.

Why Data Analysts Need Statistics

Statistics protects you from three classic mistakes:

1. Patterns in noise - seeing a "trend" that is actually random fluctuation. 2. Samples that lie - concluding something from a biased sample. 3. Overconfident numbers - reporting a precise average while ignoring uncertainty.

First Example in Python

import numpy as np

Simulate 500 order values (mix of two customer segments)

rng = np.random.default_rng(42) segment1 = rng.normal(45, 8, 350) # regular customers segment2 = rng.normal(70, 6, 150) # premium customers orders = np.concatenate([segment1, segment2])

print(f"count: {len(orders)}") print(f"mean: {orders.mean():.2f}") print(f"median: {np.median(orders):.2f}") print(f"std: {orders.std():.2f}")

Output:

count:  500
mean:   52.50
median: 48.20
std:    12.90

Notice that the mean is higher than the median. That is a clue that the distribution is right-skewed - a topic for the central tendency lesson.

Key Terminology

TermMeaning
PopulationThe entire group you want to learn about
SampleA subset of the population you actually measure
VariableA characteristic that can take different values
ParameterA number describing a population (usually unknown)
StatisticA number describing a sample (what you compute)
DistributionHow values of a variable are spread out

Common Pitfalls

  • Confusing correlation with causation - ice cream sales and drowning both rise in summer, but one does not cause the other.
  • Ignoring uncertainty - a 2% difference in conversion may vanish when you account for sample size.
  • Bias in collection - surveying only your most active users will skew every result.

Summary

  • Statistics has two branches: descriptive (summarizing) and inferential (generalizing).
  • Every analysis starts with understanding your data's distribution.
  • Statistics guards against false patterns, biased samples, and overconfidence.
  • Python (numpy, scipy, pandas) gives you a full statistics toolbox.

Next Lesson

In the next lesson, you will learn how to classify data correctly: types of data and measurement scales. This is the foundation for choosing the right statistical method later.

Quiz - Quiz - What is Statistics

1. What is the difference between descriptive and inferential statistics?

2. A number that describes a population is called a:

3. Which of these is descriptive statistics?

4. Using 500 orders to estimate the average of all future customers is an example of:

5. Why does statistics protect an analyst from overconfidence? multiple answers