Two datasets can have the identical average and be nothing alike. Eight people averaging $50,000 and eight averaging $12,500 both have a mean of $50,000 — and the second group is spread across a range the first never approaches. The standard deviation is the number that tells the two apart.
What it measures
Standard deviation is the typical distance from the mean. The calculation squares each deviation, averages them, and takes a square root:
σ = √[ Σ(x − x̄)² ÷ N ] (population)
s = √[ Σ(x − x̄)² ÷ (N − 1) ] (sample)
Two steps deserve justification. Squaring is what makes the deviations addable: without it, deviations below the mean would cancel those above and every dataset would have a spread of zero. Taking the root at the end returns the answer to the units you started in — variance lives in squared units, which is why it reads oddly next to a mean.
Population or sample, and why it matters
The only difference is the divisor: N or N−1. And the N−1 is not a rounding convention — it corrects a real bias.
Think of a factory. The true population is every part the line ever makes. You inspect ten and compute the mean. That sample mean will, on average, sit slightly below the true mean — because you are unlikely to have drawn the ten most extreme parts. The deviations you measured are therefore slightly smaller than the true spread, and dividing by N understates it. Dividing by N−1 corrects exactly that.
The rule: if you have every member, use N. If you have a subset, use N−1. When the sample is small relative to the population, the difference matters a great deal — with n = 5, using N understates the standard deviation by about 10%.
Worked example by hand
Take the classic set: 2, 4, 4, 4, 5, 5, 7, 9. Eight values, so the mean is (2+4+4+4+5+5+7+9)/8 = 40/8 = 5.
Deviations and their squares:
- 2 → −3 → 9
- 4 → −1 → 1
- 4 → −1 → 1
- 4 → −1 → 1
- 5 → 0 → 0
- 5 → 0 → 0
- 7 → +2 → 4
- 9 → +4 → 16
The squared deviations sum to 32.
As a population: variance = 32 ÷ 8 = 4, so σ = √4 = 2.
As a sample: variance = 32 ÷ 7 ≈ 4.571, so s ≈ 2.138.
Interpretation: values typically sit about 2 units from the average. The 9 sits at z = (9 − 5)/2 = +2 — two standard deviations out, and the obvious candidate for an outlier in this dataset. The standard deviation calculator does all of this and reports each figure with its unit.
The 68-95-99.7 rule
For data roughly bell-shaped, about 68% of values lie within 1σ, 95% within 2σ and 99.7% within 3σ. This is the empirical rule, and it is the practical reason the standard deviation is such a natural unit of measurement.
Quality control works this way. If a process has a mean of 50 and a standard deviation of 2, the ±3σ control limits are 44 to 56. A reading outside those is not bad luck — at that rate you would see fewer than three such in a thousand, so investigating is not optional. It is also the basis of "how many standard deviations from the mean" reporting, and of the ±2 in scientific papers.
Z-scores make datasets comparable
Two measurements on different scales cannot be compared raw. Standardising fixes that:
z = (x − mean) ÷ standard deviation
Exam scores of 88 and 72 in a class with mean 75 and s = 8 give z = +1.625 and −0.375. Battery life of 6 hours in a fleet with mean 5 and s = 0.4 gives z = +2.5. Now they are comparable: that battery is a much bigger outlier than the exam score, even though the numbers do not look related.
The convention in most fields is that |z| beyond 2 is worth a look, and beyond 3 is an outlier. With a small dataset, expect a few of these by chance — normal data produces roughly 5% of points beyond 2σ, so flagging is a prompt to investigate, not a verdict.
Coefficient of variation: comparing across scales
Standard deviation in absolute terms is hard to compare across differently sized quantities. The coefficient of variation, CV = σ ÷ mean, expresses spread relative to size, and it can be compared freely. A process with mean 100 and σ 5 has CV 5%; one with mean 10 and σ 2 has CV 20% — the second is far less consistent despite the smaller absolute deviation. In finance, CV is the risk per unit of return, which is why it sits beside Sharpe ratios.
What standard deviation does not tell you
It summarises spread, and spread is not the whole story. A symmetric, bell-shaped dataset can hide a bimodal one perfectly well; a mean of 5 with σ of 2 might describe a tidy distribution or two separate clusters, and no single number distinguishes them. That is what a histogram and a median are for. Also, σ is pulled by outliers more than most measures — a single wild value moves it considerably — so for small or skewed datasets the median and interquartile range often say more.
Coefficient of variation: comparing spread across scales
An absolute standard deviation cannot be compared between quantities of different size. Dividing by the mean fixes that:
CV = σ ÷ mean
A process with mean 100 and σ of 5 has CV = 5%. Another with mean 10 and σ of 2 has CV = 20%. The second is far less consistent despite the smaller absolute deviation, and the CV is what shows it. In finance the same ratio is risk per unit of return, which is why it appears beside Sharpe ratios; in manufacturing it is the cleanest way to compare the consistency of a 5 kg part against a 5 tonne part.
When standard deviation is the wrong tool
σ is pulled hard by outliers — one wild value moves it considerably more than it moves the median — and it assumes a roughly symmetric distribution. Where either fails, better alternatives exist:
- Interquartile range and the median, for skewed data or small samples with outliers.
- Range (max − min), which is crude but honest for very small sets and for tolerancing work.
- Mean absolute deviation, the mean's own idea of spread, less outlier-sensitive than σ and easier to explain.
The general principle: if the distribution is not roughly normal, or n is small, the robust measures are the safer summary. A p-value or a control limit computed from a badly skewed distribution is not a small error — it is a wrong number.
Weighted standard deviation
For a combined population, the variances combine in a specific way:
σ² = Σ(wᵢσᵢ²) + Σwᵢ(μᵢ − μ)²
The second term is the spread between groups, the first the spread within them. The practical consequence is Simpson's paradox: adding a subgroup can change the overall trend because it changes how much variation is between groups rather than within. Any aggregate statistic, standard deviation included, can move in a direction opposite to every subgroup it contains.
Frequently asked questions
What does standard deviation tell me?
It measures how far values typically sit from the mean. A small standard deviation means the values cluster tightly around the average; a large one means they are widely scattered. Two datasets can share a mean and have very different standard deviations.
Population or sample — which should I use?
Use population (divide by N) when your list covers the entire group you care about. Use sample (divide by N−1) when it is a subset drawn from a larger population, such as a survey.
Why is the variance the standard deviation squared?
Squaring removes the negative signs from deviations below the mean, which is necessary because a spread cannot be negative. Taking the square root returns the answer to the original units.
What is a z-score and what is it for?
A z-score measures how many standard deviations a value sits from the mean: z = (x − mean) ÷ standard deviation. Values beyond |2| are commonly flagged as unusual, which is how quality control limits and anomaly detection work.