Fraud detection. A bank’s model flags \(99\%\) of fraudulent transactions and \(1\%\) of legitimate ones. What is the chance that a flagged transaction is fraud? The answer depends on the base rate \(\hat{p}(y)\): the share of all transactions that are fraud.
Let \(\hat{x}_{i}\) be whether the model flags transaction \(i\) and \(\hat{y}_{i}\) be whether the transaction is fraud. Suppose fraud is rare, \(100\) of \(100{,}000\) transactions, so \(\hat{p}(\text{fraud}) = 0.001\). The counts are
\[\begin{array}{c|cc|c}
& \text{legitimate} & \text{fraud} & \text{Row total}\\
\hline
\text{not flagged} & 98901 & 1 & 98902\\
\text{flagged} & 999 & 99 & 1098\\
\hline
\text{Column total} & 99900 & 100 & 100000
\end{array}\]
Before seeing the model’s output, the chance that a transaction is fraud is the base rate, \(\hat{p}(\text{fraud}) = 100/100000 = 0.001\). After seeing a flag, it is the conditional, \(\hat{p}(\text{fraud}\mid\text{flagged}) = 99/1098 \approx 0.09\). The flag matters: it raises the chance of fraud \(90\)-fold. But \(91\%\) of flags are still false alarms, because the \(1\%\) false alarm rate applies to \(99{,}900\) legitimate transactions, which swamps the \(99\) true detections.
Now suppose fraud is common, \(10{,}000\) of \(100{,}000\) transactions, so \(\hat{p}(\text{fraud}) = 0.10\). The model is unchanged, but the counts become
\[\begin{array}{c|cc|c}
& \text{legitimate} & \text{fraud} & \text{Row total}\\
\hline
\text{not flagged} & 89100 & 100 & 89200\\
\text{flagged} & 900 & 9900 & 10800\\
\hline
\text{Column total} & 90000 & 10000 & 100000
\end{array}\]
and \(\hat{p}(\text{fraud}\mid\text{flagged}) = 9900/10800 \approx 0.92\). The same flag from the same model means fraud is unlikely in one setting and likely in the other. The only thing that changed is the base rate.
The base-rate fallacy is to judge the flag by the model’s accuracy alone, as if \(\hat{p}(y\mid x)\) did not depend on \(\hat{p}(y)\). Whenever the outcome is rare, even an accurate screen produces mostly false positives: medical tests for rare diseases, polygraph screening of employees, and airport security checks all share this arithmetic.
Code
# Rare fraud: rows = model output, columns = transaction type
counts_rare <- rbind(
not_flagged = c(legitimate=98901, fraud=1),
flagged = c(legitimate=999, fraud=99))
colSums(counts_rare) / sum(counts_rare) # base rate p(y)
## legitimate fraud
## 0.999 0.001
round(prop.table(counts_rare, 1), 3) # conditional p(y | x)
## legitimate fraud
## not_flagged 1.00 0.00
## flagged 0.91 0.09
# Common fraud: same model, base rate 100 times higher
counts_common <- rbind(
not_flagged = c(legitimate=89100, fraud=100),
flagged = c(legitimate=900, fraud=9900))
colSums(counts_common) / sum(counts_common)
## legitimate fraud
## 0.9 0.1
round(prop.table(counts_common, 1), 3)
## legitimate fraud
## not_flagged 0.999 0.001
## flagged 0.083 0.917