Probability is the one branch of maths that everybody already uses and nobody admits to. Every "probably", every odds quote, every weather forecast is a probability claim in disguise. The trouble is not that the arithmetic is hard — most real problems need one formula and a couple of multiplications. The trouble is that the questions are easy to ask badly and the answers are easy to over-read.
Conditional probability: the workhorse
A plain probability P(A) asks how often A happens across many trials. A conditional probability P(A|B) asks how often A happens given that you already know B. The two differ, and the size of the difference is where most reasoning errors live.
The general form is:
P(A|B) = P(B|A) × P(A) ÷ P(B)
Read it right to left. Take the event A you care about, weight it by how often you would also have seen B in that case, and divide by how often B occurs at all. The probability calculator evaluates conditional probability directly, which is worth using once you have the formula memorised purely to check your own arithmetic.
A concrete medical example, because this is the situation the formula was designed for. Suppose a condition affects 1% of people. A test is 90% sensitive, meaning it correctly flags 90% of those who have it, and 95% specific, meaning it correctly clears 95% of those who do not. Asked "if someone tests positive, do they have the condition?", intuition says 90%. The correct answer is much lower.
Work in a population of 10,000 people, which makes every count an integer:
- 100 have the condition. Of those, 90 test positive.
- 9,900 do not. Of those, 5% test false-positive: 495.
- Total positive tests: 90 + 495 = 585.
So P(condition | positive) = 90/585 = 0.1538, about 15.4% — not 90%. The sensitivity describes the test's behaviour on people who are sick, and says nothing about how common being sick is. This is base rate neglect, and it is the single most consequential probability error in everyday life: the rarer the condition, the more the false positives swamp the true ones, no matter how good the test. A test that is 99% accurate on a disease with 0.01% prevalence produces mostly false alarms.
Independence versus mutual exclusion
These two get confused constantly, and they are opposites rather than relatives.
Mutually exclusive means the events cannot both happen: drawing a heart and drawing a spade in one card. Their probabilities add, P(A and B) = 0.
Independent means one event gives you no information about the other: two separate coin flips, or a person's height and their favourite colour. Both can happen together freely, and P(A and B) = P(A) × P(B).
The distinction bites in sampling. Drawing from a deck without replacement creates a dependence between draws, which is where the classic card example goes wrong.
Worked example: two cards of the same suit. Draw two cards from a standard 52-card deck without replacing the first. The first card is some suit; of the remaining 51 cards, 12 share that suit. So the probability is 12/51 = 4/17 ≈ 0.2353.
It is tempting to write (13/52)² = 1/16 = 6.25% instead, and that is a common error worth understanding rather than just avoiding. That expression treats the second draw as independent of the first — as if the first card had been put back, restoring all 52. The independence assumption is what fails: knowing the first card's suit reduces the pool. The fraction converter is a quick way to keep the exact and decimal forms side by side.
The birthday problem: 23 people and a coin flip
Ask 23 people their birthdays, ignoring years. The chance that at least two share a date is 0.5073 — a little better than even. Most people find this implausible, and the reason is that the events feel independent when they are not. The second person's birthday is *nearly* unconstrained; by the twentieth it is clearly constrained.
Compute it as a chain. The probability that all birthdays are distinct is
365/365 × 364/365 × 363/365 × ... × 343/365
Multiplying those 23 factors gives approximately 0.4927 of the time with no match, so the probability of at least one shared birthday is 1 − 0.4927 = 0.5073.
The general rule: with n people the probability of a shared birthday approaches 50% at n = 23, and rises slowly thereafter — 30 people gives 0.7063, and 40 people gives about 0.891. The curve is surprisingly flat near the middle and then climbs steeply. For intuition on why the factors shrink, note that the 23rd person must avoid 22 already-taken dates out of 365, and by the 40th they must avoid 39.
Monty Hall: why switching doubles your chance
Three doors, a car behind one, goats behind the others. You pick door 1. The host, who knows where the car is and always opens a door with a goat, opens door 3 to reveal a goat. Should you switch to door 2?
Switching wins 2/3 of the time. The reason is that the host's action is not random, and treating it as though it were is the entire source of the error.
Enumerate your first pick:
- Car behind door 1 (probability 1/3): staying wins, switching loses.
- Car behind door 2 (probability 1/3): the host is forced to open door 3, so switching wins.
- Car behind door 3 (probability 1/3): the host is forced to open door 2, so switching wins.
Switching collects the car in two of three equally likely cases: 2/3. Your original 1/3 guess never improved, but the host's deliberate action concentrated the remaining 2/3 into the single door you did not pick. Switching cannot help if your first pick was right, and cannot hurt if it was wrong — which is exactly what makes it dominant.
Expected value: an average, not a prediction
Expected value = Σ (outcome × probability). It is a weighted average across many repetitions. What it is not is a forecast of any single outcome, and that gap is the most useful thing to understand about probability.
A fair raffle ticket costing 1, with a 1-in-1,000 chance of paying 1,000, has expected value (999/1,000) × 0 + (1/1,000) × 1,000 = 1.00. Break-even, fair. Yet the outcome is binary: you either win 1,000 or lose 1, and the most likely single outcome by far is losing. Expected value sitting exactly between the two possibilities tells you nothing about which one occurs.
Expected value can even fall outside the range of possible single outcomes. Choose a number from −1 to 6 with equal probability and the expected value is 2.5 — not a number you can draw. Every game with a negative expected value is priced off this concept: a lottery operator paying out 1,000 in 1,000 draws on a 1 ticket takes in 1,000 in ticket sales and pays out 1, so the operator's expected profit per ticket is +0.999. That is not a prediction of profit on any particular ticket, and it is the mechanism regardless of how many people play. Insurance works the same way in reverse: the expected payout is less than the premium, pooled across many low-probability events, so it is a bad bet one at a time and a good one across a portfolio. The standard deviation calculator is the natural companion here, since expected value and spread are the two numbers that describe a distribution — and the standard deviation guide covers how spread is defined.
Sampling, and why a biased sample beats a large one
A large random sample is good. A large biased sample is worse than a small random one, because the bias does not shrink with size — it just becomes more confidently wrong. Every source of bias has the same shape: the sampling mechanism correlates with the thing being measured.
Asking "how do you feel about the economy?" at a stock exchange is not a small-sample problem, it is a category error. Polling only people who answer a landline misses the younger and poorer ones entirely. Quitting a survey after the easy questions selects for people with strong opinions. Each of these produces a stable, repeatable, entirely fictional number, which is why confident polling can be reliably wrong.
Risk versus uncertainty
Economists draw a sharp line here. Risk means you know the possible outcomes and their probabilities — an actuarial table, a shuffle of a deck, a loaded die you have measured. Uncertainty means you do not, and no amount of calculation will produce a number. We do not have a reliable probability distribution for the outcome of the next war, and anyone quoting one is doing something other than probability.
The practical significance is that risk can be managed with arithmetic and uncertainty cannot. Hedging a known risk is a calculation; hedging an unmeasurable one is a guess dressed as a calculation. When someone presents a single precise percentage for something genuinely uncertain, the precision is the tell.
Things that look like probabilities and are not
Forecast percentages do not add. A 30% chance of rain on each of three days is not a 90% chance over the weekend. Assuming independence, the chance of no rain on all three is 0.7 × 0.7 × 0.7 = 0.343, so the chance of rain on at least one day is 1 − 0.343 = 0.657, not 0.90. Summing forecast percentages is one of the most widespread errors in everyday numeracy, and the percentage calculator will convert between the two forms while you check it.
"Odds" and "probability" are not interchangeable. Odds of 3 to 1 mean p = 3/4, while a "3-in-4 chance" means the same thing — but "odds of 3 to 1" in a betting context means the bettor receives 3 times the stake in profit if they win, which already has the house margin baked into it. Converting between the two is one of the most common slips in sports reporting.
Coin flips are not really fair. Real coins are slightly biased by manufacturing asymmetry, and "random" digit sources in computers are pseudorandom. For practical maths these are negligible, but the honest statement is that the 50% assumption is a model rather than a fact about physics.
Frequently asked questions
What is conditional probability and when do I use it?
Conditional probability is the probability of one event given that another has occurred, written P(A|B). It differs from the plain probability P(A) because knowing B changes what is plausible. You need it whenever a rare event is being tested for and the result is more common among people who have the condition than among those who do not.
Why is expected value not the most likely outcome?
Expected value is a long-run average across many repetitions, not a prediction about the next event. It can even fall outside the range of possible single outcomes. This is why a fair lottery ticket has a positive expected value for the operator and a negative one for you, and why average insurance payouts do not describe any individual policy.
Does independence mean the events cannot happen together?
No — that describes mutually exclusive events, which is the opposite idea. Independent events can occur together freely; knowing one occurred gives no information about the other. Two coin flips on a fair coin are independent, while drawing two aces from a deck without replacement creates a dependence between the two draws.
Why does the Monty Hall problem defeat most people?
Because the host's action is not random and everyone treats it as if it were. Switching wins 2 of the time because, in the two cases where your first pick was wrong, you are guaranteed to be offered the remaining prize — so switching collects the prize twice, while staying collects it once.