An alert has two possible routes: a fault triggered it, or a healthy machine triggered it. Bayesian inference compares the probability mass arriving by each route.
The key move: after seeing an alert, restrict attention to machines with an alert. The fraction that are faulty is the posterior probability.
P(fault | alert) = true alertstrue alerts + false alerts
The original prediction, worked through
- Out of 10,000 machines, 100 have a fault and 9,900 are healthy.
- 90% of 100 gives 90 true alerts. The remaining 10 faults have no alert.
- 5% of 9,900 gives 495 false alerts. The other 9,405 healthy machines have no alert.
- Among the 585 alerts, 90 are true: 90 / 585 = 2 / 13 ≈ 15.38%.
The alert is still useful evidence: it raises the probability from 1% to about 15%. But 90% sensitivity answers a different question. Its denominator is faulty machines; the posterior's denominator is alerted machines.
This uses conditional probability and the law of total probability. See MIT 18.05, reading 3.
Open the mathematics: from counts to Bayes' rule
- H
- The hypothesis: this machine has a fault. ¬H means no fault.
- +
- The observed evidence: the detector raises an alert.
- p = P(H)
- Prior probability, here the relevant population's base rate.
- s = P(+ | H)
- Sensitivity: probability of an alert if there is a fault.
- f = P(+ | ¬H)
- False-positive rate: probability of an alert if there is no fault.
The vertical bar means “given.” In formulas, use probabilities from 0 to 1: 5% is 0.05.
P(H | +) = s × ps × p + f × (1 − p)
The numerator is P(H and +) = P(+ | H)P(H). The denominator adds both mutually exclusive ways an alert can occur: P(+) = sp + f(1 − p). Dividing their intersection by P(+) is the definition of conditioning.
For a population of size N, expected true alerts are Nsp and false alerts are Nf(1 − p). N cancels in their ratio. A larger fleet changes counts, but not this probability if the rates stay fixed.
Odds make evidence strength explicit
For 0 < p < 1 and f > 0:
posterior odds = prior odds × likelihood ratio
odds = probability / (1 − probability)
likelihood ratio for an alert = s / f
In the first example, prior odds are 1 / 99 and the likelihood ratio is 0.9 / 0.05 = 18. Posterior odds are 18 / 99 = 2 / 11. Converting odds to probability gives 2 / (2 + 11) = 2 / 13.
An alert supports the fault hypothesis when s > f, leaves it unchanged when s = f, and counts against it when s < f, provided the prior is not certain and the observation is possible.
What about no alert?
P(H | no alert) = (1 − s) × p(1 − s) × p + (1 − f) × (1 − p)
At the starting settings this is 10 / 9,415 ≈ 0.1062%. The chance of a healthy machine after no alert is its complement, about 99.89%.
When a result is undefined
If sp + f(1 − p) = 0, the model predicts no alerts at all. Conditioning on an alert is undefined, not 0%. The same applies to “no alert” when its probability is zero. Try these boundaries in the experiment. A supposedly impossible observation is a reason to revisit the model.
What this model assumes
The two states are exhaustive and mutually exclusive. The base rate applies to the population being examined, and the sensitivity and false-positive rate apply there too. We treat the rates as known and the true fault state as well defined.
Real rates are estimated with uncertainty and can change with operating conditions, detector thresholds, or how machines are selected. A single binary observation needs no independence assumption. Multiplying likelihood ratios across observations requires the right conditional model. Repeated evidence is not automatically independent.