CCalclabhub

Statistics

How To Calculate Standard Deviation

Standard deviation is the typical distance of observations from their mean, measured in the original units. The squaring is what makes it algebraically useful — and the n-1 is what makes it honest for a sample.

Quick Answer

SD = sqrt(Sum of (xi - Mean)^2 / (n - 1))

xi
Each observation
mu or x-bar
Mean of the observations
n
Number of observations
Sigma
Sum of the squared deviations from the mean

Subtract the mean from each value, square those deviations, sum them, divide by one less than the count for a sample, then take the square root. For the data set 10, 12, 23, 23, 16 the mean is 16.8 and the sum of squared deviations is 146.8, giving a sample standard deviation of 6.0581 and a population standard deviation of 5.4185. Because the result is in the same units as the data, it can be read directly as a typical distance from the average.

What Is Standard Deviation?

Two data sets can share a mean and look completely different. Ten values clustered tightly around fifty and ten values scattered from five to ninety-five both average to around fifty, and no average will distinguish them. Standard deviation fills exactly that gap: it is a single number describing how far values typically sit from the centre, expressed in the same units as the measurements themselves.

The construction has three moves, and each one solves a specific problem. First subtract the mean, giving signed distances. Then square them. Then average and take the square root at the end. The squaring exists because summing signed deviations always gives exactly zero by definition — the positive and negative distances cancel — so some way of removing the signs is required before averaging means anything.

Squaring is not the only option and it is worth knowing why it won. Taking absolute values would be more literal and would give the mean absolute deviation, but absolute values are algebraically awkward: the absolute value function has a kink at zero that breaks the clean calculus underlying almost every statistical procedure. Squares are smooth, differentiable, and lead naturally to the variance-additivity results that make statistical inference possible.

The cost is interpretability, and it is the reason the square root comes back at the end. Squared deviations are measured in squared units — dollars squared, seconds squared — which nobody can reason about. Dividing by the count gives the variance, still in squared units, and taking the square root returns values to the original measurement scale. Variance is the mathematically convenient quantity; standard deviation is the human-readable one.

Then there is the divisor, which is where genuine subtlety lives. Dividing by n describes the data you actually have, treating it as the entire population of interest. Dividing by n-1 estimates the spread of the larger population you drew the sample from. The intuition: any sample mean is pulled slightly toward the sample's own values, making it closer to them than the true population mean would be. Squared deviations around the sample mean therefore come out systematically too small, and replacing n with n-1 removes almost exactly that bias.

The size of the correction is anything but academic at small counts. Dividing by 5 instead of 4 here gives a variance of 29.36 rather than 36.70 — a twenty percent difference, which is why sample versus population is not a rounding question. As n grows the two converge: at n = 100 the difference is one percent, and with large samples the choice barely matters. Below about thirty observations it matters a great deal.

Reading the number requires knowing the shape. If the distribution is roughly bell-shaped, about 68% of observations fall within one standard deviation of the mean and about 95% within two — the familiar empirical rule. For our data that means roughly 68% of values sit between 10.74 and 22.86. That interpretation depends heavily on symmetry, and for skewed or heavy-tailed distributions the rule can be badly wrong even though the calculation is identical.

This sensitivity explains standard deviation's main weakness: it treats a value far above the mean and one equally far below as identical contributions, which is misleading for one-sided risks. Investment analysts distinguish downside deviation for exactly this reason, and any situation where losses and gains are not symmetric deserves the same scrutiny.

Finally, the units are the feature that makes it practical. A standard deviation of 6.06 for data measured in minutes means about six minutes of typical spread, directly comparable to the mean of 16.8 minutes. The coefficient of variation — standard deviation divided by the mean, here 36.06% — removes units entirely and allows comparison between quantities that could never otherwise be put side by side.

Formula

s = sqrt(Sum(xi - Mean)^2 / (n - 1))

Use when the data is a sample drawn from a larger population. Dividing by n-1 removes the downward bias in estimating that population's spread.

SymbolMeaning
sSample standard deviation
x-barSample mean
nNumber of observations

sigma = sqrt(Sum(xi - mu)^2 / n)

Use when the values are the entire population of interest rather than a sample from it — every student in the class, every item in the batch.

SymbolMeaning
sigmaPopulation standard deviation
muPopulation mean

CV = Standard Deviation / Mean x 100

Spread expressed as a percentage of the centre. Unitless, so it can compare variability across quantities measured in different units.

SymbolMeaning
CVCoefficient of variation

How To Calculate Standard Deviation

  1. 1

    Compute the mean

    Add every value and divide by how many there are: (10 + 12 + 23 + 23 + 16) / 5 = 16.8. Everything downstream depends on this being right.

  2. 2

    Find each deviation from the mean

    Subtract the mean from every observation: -6.8, -4.8, 6.2, 6.2 and -0.8. These sum to zero by construction, which is a convenient check before continuing.

  3. 3

    Square each deviation

    46.24, 23.04, 38.44, 38.44 and 0.64. Squaring removes the signs so distances of equal size on either side contribute equally.

  4. 4

    Sum them and choose the divisor

    Total is 146.8. Divide by n-1 = 4 for a sample giving 36.70, or by n = 5 for the whole population giving 29.36. This is the variance.

  5. 5

    Take the square root

    sqrt(36.70) = 6.0581. Bringing the units back to the original scale makes the figure readable as roughly six units of typical distance from the mean of 16.8.

Examples

Example 1: Five observations worked through

Data
10, 12, 23, 23, 16
Count
5
StepCalculationResult
Mean(10 + 12 + 23 + 23 + 16) ÷ 516.8
Deviations from the mean10 - 16.8, 12 - 16.8, 23 - 16.8, 23 - 16.8, 16 - 16.8-6.8, -4.8, 6.2, 6.2, -0.8
Squared deviations(-6.8)^2 + (-4.8)^2 + (6.2)^2 + (6.2)^2 + (-0.8)^2146.8 in total
Sample variance146.8 ÷ 436.70
Sample standard deviationsqrt(36.70)6.0581

Result: 6.0581 — the same data treated as a population would give sqrt(146.8 ÷ 5) = 5.4185, about twenty percent smaller.

Example 2: How much the divisor choice costs

Sum of squared deviations
146.8
n
5
StepCalculationResult
Dividing by n for the population146.8 ÷ 529.36 variance
Population standard deviationsqrt(29.36)5.4185
Dividing by n - 1 for a sample146.8 ÷ 436.70 variance
Sample standard deviationsqrt(36.70)6.0581
Relative difference between the two answers6.0581 ÷ 5.4185 - 111.80% higher

Result: 5.4185 as a population, 6.0581 as a sample — the sample figure is 11.80% larger because the smaller divisor corrects for using an estimated mean.

Example 3: Reading the result

Mean
16.8
Sample SD
6.0581
StepCalculationResult
Within one standard deviation16.8 - 6.0581 and 16.8 + 6.058110.74 to 22.86
How much of the data sits in that bandthree of five observations60.00%, against roughly 68% for a bell-shaped distribution
Range for comparison23 - 1013.0
Coefficient of variation6.0581 ÷ 16.836.06%

Result: 10.74 to 22.86 holds 60.00% of these five observations, and the spread is 36.06% of the mean — comparable across units in a way the raw number is not.

Calculator

Standard deviation

6.0581

Variance
36.7
Mean
16.8
Spread relative to the mean
36.06%

Values update as you type. This calculator covers the single scenario its formula assumes — see Common Mistakes for what it leaves out.

Prefer a full-width tool? Open the Standard Deviation calculator page.

Common Mistakes

  • Forgetting to take the square root

    Reporting the variance as if it were the standard deviation inflates the figure and puts it in squared units. A variance of 36.70 is a standard deviation of 6.0581.

  • Dividing by n when the data is a sample

    It systematically understates population spread, because deviations are measured around a sample mean that was itself fitted to those same points. Use n-1 unless you truly have every member.

  • Averaging the absolute deviations instead of squaring

    That computes mean absolute deviation, a legitimate but different statistic. It is smaller and does not share the inferential properties of the variance.

  • Applying the 68-95 rule without checking the shape

    Those percentages assume approximate normality. Skewed or heavy-tailed data can put far more or far less inside one standard deviation while the arithmetic is identical.

  • Comparing spread across different scales using raw values

    A standard deviation of 6 minutes against one of 6 dollars says nothing about which is more variable. Divide by the mean to compare, and avoid doing so when means sit near zero.

FAQ

When should I use n and when n-1?

Use n-1 whenever your data is a sample drawn from something larger, which is nearly always in practice. Use n only when the observations genuinely are the entire population of interest, such as every item in a production batch.

Why square the deviations instead of taking absolute values?

Squared deviations are smooth and differentiable, which makes the underlying mathematics tractable and gives variance useful additive properties. Absolute deviations are more literal but algebraically clumsy, and produce a different statistic rather than an equivalent one.

What units is standard deviation measured in?

The same units as your data, because the final square root undoes the squaring. Variance, by contrast, is in squared units and is mainly useful as an intermediate quantity.

Can standard deviation be larger than the mean?

Yes, and it is common for skewed data with small or near-zero means, including anything that can be negative. It simply indicates that typical observations sit further from the centre than the centre is from zero.

How many observations do I need for a useful standard deviation?

Mechanically two suffice, though they give enormous estimation error. For practical purposes thirty or more makes the estimate reasonably stable, and the n versus n-1 distinction negligible, though you should still use the correct convention.

References

  1. [1]Khan Academy, Standard deviation and variance — measures of spread — https://www.khanacademy.org/math/statistics-probability/summarizing-quantitative-data
  2. [2]National Institute of Standards and Technology (NIST), Engineering Statistics Handbook: measures of scale and variability — https://www.itl.nist.gov/div898/handbook/eda/section3/eda35.htm
  3. [3]Wolfram MathWorld, Bessel's correction and unbiased estimation of variance — https://mathworld.wolfram.com/BesselsCorrection.html