Bayesian inference

An alert is evidence.
How much evidence?

A detector catches almost every fault. Yet most of its alerts can still be false. Explore the missing piece: how common the fault was before the alert.

Make a prediction.

A machine raises an alert.

In a hypothetical fleet, 1% of machines have a fault. The detector alerts on 90% of faulty machines and 5% of healthy machines. A randomly chosen machine has an alert.

What is the chance it actually has a fault?

1%base rate
90%sensitivity
5%false-positive rate

All numbers in this lesson are invented teaching examples. No fleet or detector was measured.

Commit to an answer before opening the experiment.

You will learn to distinguish a detector's sensitivity from the probability of a fault after an alert, explain the denominator, and recognize when the model's assumptions fail. Percentages are enough to begin.

Change the world. Read the evidence.

The base rate describes the machines before checking alerts. Sensitivity is the fraction of faulty machines that alert. The false-positive rate is the fraction of healthy machines that alert.

Start with your prediction above, or open the experiment now.

Change only the base rate

Predict what happens at 20%, then try it. Why does an unchanged detector now produce mostly true alerts?

Improve one thing

Start at 1%, 90%, 5%. Compare raising sensitivity to 100% with lowering false positives to 0.5%. Which change makes an alert more trustworthy here?

Make evidence useless

Set sensitivity equal to the false-positive rate, with a rate above 0%. Does an alert change the prior? Explain why.

Check the experiment predictions

At 20% base rate, 1,800 of 2,200 alerts are true: 81.82%. The healthy group is smaller and the faulty group larger.

At 1% base rate, raising sensitivity to 100% gives 100 / 595 = 16.81%. Lowering the false-positive rate to 0.5% with sensitivity at 90% gives 90 / 139.5 = 64.52%. Here, the large healthy group makes false alarms the dominant issue. This does not imply detector thresholds can improve sensitivity and false positives independently in practice.

If sensitivity equals the false-positive rate, an alert is equally likely in either state. For an alert with nonzero probability, the posterior equals the prior.

Count the explanations for the alert.

An alert has two possible routes: a fault triggered it, or a healthy machine triggered it. Bayesian inference compares the probability mass arriving by each route.

The key move: after seeing an alert, restrict attention to machines with an alert. The fraction that are faulty is the posterior probability.

P(fault | alert) =

The original prediction, worked through

  1. Out of 10,000 machines, 100 have a fault and 9,900 are healthy.
  2. 90% of 100 gives 90 true alerts. The remaining 10 faults have no alert.
  3. 5% of 9,900 gives 495 false alerts. The other 9,405 healthy machines have no alert.
  4. Among the 585 alerts, 90 are true: 90 / 585 = 2 / 13 ≈ 15.38%.

The alert is still useful evidence: it raises the probability from 1% to about 15%. But 90% sensitivity answers a different question. Its denominator is faulty machines; the posterior's denominator is alerted machines.

This uses conditional probability and the law of total probability. See MIT 18.05, reading 3.

Open the mathematics: from counts to Bayes' rule
H
The hypothesis: this machine has a fault. ¬H means no fault.
+
The observed evidence: the detector raises an alert.
p = P(H)
Prior probability, here the relevant population's base rate.
s = P(+ | H)
Sensitivity: probability of an alert if there is a fault.
f = P(+ | ¬H)
False-positive rate: probability of an alert if there is no fault.

The vertical bar means “given.” In formulas, use probabilities from 0 to 1: 5% is 0.05.

P(H | +) =

The numerator is P(H and +) = P(+ | H)P(H). The denominator adds both mutually exclusive ways an alert can occur: P(+) = sp + f(1 − p). Dividing their intersection by P(+) is the definition of conditioning.

For a population of size N, expected true alerts are Nsp and false alerts are Nf(1 − p). N cancels in their ratio. A larger fleet changes counts, but not this probability if the rates stay fixed.

Odds make evidence strength explicit

For 0 < p < 1 and f > 0:

posterior odds = prior odds × likelihood ratio
odds = probability / (1 − probability)
likelihood ratio for an alert = s / f

In the first example, prior odds are 1 / 99 and the likelihood ratio is 0.9 / 0.05 = 18. Posterior odds are 18 / 99 = 2 / 11. Converting odds to probability gives 2 / (2 + 11) = 2 / 13.

An alert supports the fault hypothesis when s > f, leaves it unchanged when s = f, and counts against it when s < f, provided the prior is not certain and the observation is possible.

What about no alert?

P(H | no alert) =

At the starting settings this is 10 / 9,415 ≈ 0.1062%. The chance of a healthy machine after no alert is its complement, about 99.89%.

When a result is undefined

If sp + f(1 − p) = 0, the model predicts no alerts at all. Conditioning on an alert is undefined, not 0%. The same applies to “no alert” when its probability is zero. Try these boundaries in the experiment. A supposedly impossible observation is a reason to revisit the model.

What this model assumes

The two states are exhaustive and mutually exclusive. The base rate applies to the population being examined, and the sensitivity and false-positive rate apply there too. We treat the rates as known and the true fault state as well defined.

Real rates are estimated with uncertainty and can change with operating conditions, detector thresholds, or how machines are selected. A single binary observation needs no independence assumption. Multiplying likelihood ratios across observations requires the right conditional model. Repeated evidence is not automatically independent.

Recognize the same structure elsewhere.

The arithmetic transfers when you can define a hypothesis, evidence, and credible probabilities. The domain determines whether those inputs are defensible.

Classification

Spam filtering

H is “this message is spam.” Evidence could be a flagged phrase. Its prevalence in spam and legitimate mail matters, along with the prior spam rate.

Boundary: correlated words cannot be treated as independent evidence without justification. A changed inbox can invalidate old rates.

Scientific evidence

A rare astronomical signal

H is “a source contains the target phenomenon.” An instrument signature updates the prior through how likely that signature is with and without the phenomenon.

Boundary: choosing candidates after scanning many signals changes the relevant model. A posterior alone does not establish a causal explanation.

Health statistics

A screening result

Sensitivity and specificity describe conditional test behavior. Predictive value also depends on prevalence in the tested population.

Boundary: population figures alone do not determine an individual's diagnosis. Clinical context and applicability of test estimates matter. See the diagnostic-performance reference.

Belief is not a decision rule. Whether to inspect a machine depends on the costs of inspection, missed faults, and unnecessary shutdowns. Bayes' rule updates a probability; it does not choose those costs.

Close the explanation. Try somewhere new.

Answer before revealing the worked feedback. On a later day, return directly here and try without looking above. Your typed notes are temporary and are not saved or synchronized.

Recognition & explanation

Reveal reasoning and self-check

They reversed P(alert | fault) and P(fault | alert). Sensitivity uses faulty machines as its denominator. The posterior uses alerted machines. You also need the relevant base rate and false-positive rate.

Self-check: name both denominators and explain why healthy machines contribute to the alert pool. Repeating “base-rate neglect” without the mechanism is recognition, not yet explanation.

Near transfer · new quantities

Reveal worked solution

In 10,000 fragments, 200 come from the kiln and 9,800 do not. Expect 160 marked kiln fragments and 392 marked other fragments. The posterior is 160 / 552 = 20 / 69 ≈ 28.99%.

Self-check: both routes to the marker must appear in the denominator. The marker raises the prior from 2% substantially, yet the kiln is still less likely than all alternatives combined. The sampling and marker-rate assumptions are essential.

Far transfer · dependent evidence

Reveal reasoning

A copied display contains no additional evidence once the original reading is known. Squaring the ratio double-counts the same observation. The second update needs P(second evidence | H, first evidence) divided by P(second evidence | ¬H, first evidence).

For an exact copy after the first alert, both probabilities are 1, giving a ratio of 1. Multiplying the original ratio again is justified only if a new observation has those same rates and is conditionally independent of the first under both H and ¬H.

Self-check: explain the dependency and name the conditional assumption. Two devices or two reports do not by themselves establish it.

Far transfer · model limits

Reveal reasoning

The relevant prior is conditional on what is already known, including unusual vibration. You need P(fault | unusual vibration), or a joint model that updates on vibration and detector evidence without double-counting. Detector performance may also differ in that selected group.

A shutdown decision needs consequences and alternatives: damage risk, inspection cost, downtime, and what additional information can be gathered. Probability alone does not set the action threshold.

Self-check: identify the selected population and separate inference from decision-making. Do not invent a replacement rate from the information given.

Record evidence of understanding

In the private curriculum, record the date and whether each answer was unaided, with support, or needs a revisit. Save one sentence about the misconception or assumption you noticed.

Keep recognition, explanation, and transfer separate. This page cannot assess free-text answers or certify mastery. A later successful attempt with a fresh problem is stronger evidence than an answer revealed today.

Sources & scope

The mathematical model is a binary Bayesian update with fixed rates. It does not estimate those rates, simulate a random sample, or model a changing machine over time.

  1. MIT 18.05, Reading 3: Conditional Probability, Independence and Bayes' Theorem (2022). Definitions, the law of total probability, Bayes' rule, and the base-rate fallacy. The fleet, pottery, and dashboard examples here are original teaching examples.
  2. CDC: Summary of Screening Test Measures (2003), page 12. Definitions and formulas for sensitivity, specificity, predictive values, and prevalence. This historical teaching source is used for its mathematics, not current clinical guidance.

Version 1 · 2026-10-10. References need a connection; the lesson does not. Author-reviewed calculations and browser checks are recorded in the private learning registry. Learner understanding has not been assessed.